1. Home
  2. Operations & SLA
OPERATIONS & SLA

Measured per allocation unit
and per cluster. Run 24×7.

A named 24×7 NOC operated with [our operating partner], escalating to Nachster AI. Availability is measured at both the allocation-unit and the cluster level, with defined exclusions and a service-credit process.

≥99.5%Monthly availability per allocation unit
24×7Monitoring and first response
Acceptedmeans in service — billing starts at COD
SLA

What we commit to.

  • ≥99.5% monthly availability per allocation unit, plus a cluster-level availability commitment. The allocation unit is one node for B300 and one rack for the rack-scale generations; it is named in the allocation offer.
  • 24×7 monitoring, hardware fault triage, and spare parts held on site.
  • Maintenance windows are pre-notified and customer-approved; we will work around the boundaries of a training run.
  • Monthly availability reporting and a dashboard you can look at any time.
  • Assured Capacity option: Hot-spare nodes are held and a usable node count is guaranteed.

Exclusions — customer-caused faults, pre-approved maintenance, force majeure — and the service-credit calculation are set out in the SLA & Operations Sheet, issued after qualification.

The acceptance standard.

A cluster is only delivered once it has passed every test below. Billing starts on the acceptance date — not on the day the agreement was signed.

  • GPU burn-in — Sustained load across every GPU, recording thermals, clocks and correctable / uncorrectable errors.
  • NCCL all-reduce — Collective performance measured across the whole cluster and checked against the reference figure.
  • Storage throughput — Read and write bandwidth and IOPS on the shared storage, measured with the pattern of the intended workload.
  • Fabric validation — Cabling topology, rail assignment, link quality and non-blocking behaviour.
  • Monitoring handover — Confirming that metrics and alerts flow both to the NOC and to your dashboard.

The pass thresholds for each test are set in the allocation offer once the final BOM is fixed. Results are disclosed to you and attached to the acceptance certificate.

What happens when something breaks.

DETECTION

The NOC sees it first

Node, fabric, power and environmental metrics are watched continuously. We do not wait for you to call.

FIRST RESPONSE

Triaged on site

Handled to the point of node replacement with spares held on site. The failed node is taken out of the cluster and a spare goes in.

ESCALATION

To a named contact

Anything first response cannot close comes to Nachster AI. Named individuals and the escalation path are fixed at contracting and can be disclosed to your auditors.

NEXT STEP

Tell us what you need. We reply within 48 hours.

Qualified enquiries receive proposed times for a technical meeting within 48 hours, and an allocation offer after that meeting. Information is handled under NDA.