- Effective
- On publication
- Sections
- 10
- Applies to
- Reserved clusters of 64 GPUs or more, held for 1 month or longer
This document lists what Fantasti commits to for every reserved cluster before billing starts and during the term. Each order form names the version it follows. A new version never changes an order form already signed.
1. Scope
These requirements apply to reserved clusters of 64 GPUs or more with a term of 1 month or longer. They do not apply to on-demand instances, spot capacity or Sandboxes.
See also
2. Fabric
All nodes in a reservation are connected by one InfiniBand fabric. Before you sign, Fantasti discloses for the pool: the fabric generation and port speed (for example NDR, 400 Gb/s per port), the number of rails per node, and the oversubscription ratio between leaf and spine.
See also
3. Tenancy
Nodes are single-tenant for the full term. No other account's workload runs on them.
4. Acceptance
Before the term starts, every node and the cluster as a whole pass:
- Burn-in of every node under sustained GPU load, for the duration stated in your order form.
- NVIDIA DCGM diagnostics at level 3 on every GPU.
- An NCCL all-reduce test across the full reservation, with bus bandwidth reported per node pair.
- InfiniBand link checks: every port at its rated speed, with error counters clean.
- A storage throughput check, where storage is part of the reservation.
Results are shared with you. Billing for the term starts when the reservation passes acceptance.
Before the term starts, every node and the cluster as a whole pass five tests: burn-in under sustained GPU load, NVIDIA DCGM diagnostics at level 3 on every GPU, an NCCL all-reduce test across the reservation, InfiniBand link checks and, where storage is reserved, a storage throughput check. The results are shared with you. Billing for the term starts when the reservation passes acceptance. Connections: Acceptance tests to Results; Results to Acceptance; Acceptance to Billing starts.
On every node, for theduration in your orderform
$ dcgmi diag -r 3On every GPU
$ all_reduce_perfAcross the fullreservation
- speed
- rated
- errors
- clean
On every InfiniBandport
- check
- throughput
Where storage is partof the reservation
- burn-in
- every node
- diagnostics
- every GPU
- bus bandwidth
- per node pair
- links
- every port
- storage
- if reserved
The reservation passes
Every node, and the cluster as awhole
On every node, for theduration in your orderform
$ dcgmi diag -r 3On every GPU
$ all_reduce_perfAcross the fullreservation
- speed
- rated
- errors
- clean
On every InfiniBand port
- check
- throughput
Where storage is part of thereservation
- burn-in
- every node
- diagnostics
- every GPU
- bus bandwidth
- per node pair
- links
- every port
- storage
- if reserved
The reservation passes
Every node, and the cluster as a whole
On every node, for the duration in yourorder form
$ dcgmi diag -r 3On every GPU
$ all_reduce_perfAcross the full reservation
- speed
- rated
- errors
- clean
On every InfiniBand port
- check
- throughput
Where storage is part of the reservation
- burn-in
- every node
- diagnostics
- every GPU
- bus bandwidth
- per node pair
- links
- every port
- storage
- if reserved
The reservation passes
Every node, and the cluster as a whole
- Billing follows acceptance
5. Your verification
After acceptance you have an acceptance window, stated in your order form, to run your own NCCL and training benchmarks. If the cluster does not meet the disclosed specification, Fantasti remediates before the term starts.
6. Node replacement
A node that fails during the term is replaced within the replacement window stated in your order form, counted from the failure being confirmed. Time a node is unavailable is credited at the reservation rate.
7. Support
Response-time targets are set by severity.
| Severity | Meaning | Response-time target |
|---|---|---|
| Severity 1 | Cluster unusable | Stated in your order form |
| Severity 2 | Degraded | Stated in your order form |
| Severity 3 | Question or minor issue | Stated in your order form |
Every reservation has a named escalation contact and a shared channel. The named contact is an AI agent. Every agent identifies itself as an AI in its first message and in its signature. These targets are separate from the uptime commitment in the SLA.
8. Disclosure
The region and facility behind the reservation are named in the order form. Where the capacity layer requires it, disclosure is under NDA.
9. Not covered
Spot capacity, on-demand instances, Sandboxes, and networking between fabrics (InfiniBand does not cross fabrics).
See also
10. Versions
Changes are published as a new version at a new address. Version history: v1, first publication.