Reserved clusters10 sections

Cluster requirements, version 1.

What every Fantasti reserved cluster commits to before billing starts: fabric, tenancy, acceptance, replacement, support and disclosure.

Version

Version 1.

Named in each order form

Effective
On publication
Sections
10
Applies to
Reserved clusters of 64 GPUs or more, held for 1 month or longer

This document lists what Fantasti commits to for every reserved cluster before billing starts and during the term. Each order form names the version it follows. A new version never changes an order form already signed.

1. Scope

These requirements apply to reserved clusters of 64 GPUs or more with a term of 1 month or longer. They do not apply to on-demand instances, spot capacity or Sandboxes.

2. Fabric

All nodes in a reservation are connected by one InfiniBand fabric. Before you sign, Fantasti discloses for the pool: the fabric generation and port speed (for example NDR, 400 Gb/s per port), the number of rails per node, and the oversubscription ratio between leaf and spine.

3. Tenancy

Nodes are single-tenant for the full term. No other account's workload runs on them.

4. Acceptance

Before the term starts, every node and the cluster as a whole pass:

  1. Burn-in of every node under sustained GPU load, for the duration stated in your order form.
  2. NVIDIA DCGM diagnostics at level 3 on every GPU.
  3. An NCCL all-reduce test across the full reservation, with bus bandwidth reported per node pair.
  4. InfiniBand link checks: every port at its rated speed, with error counters clean.
  5. A storage throughput check, where storage is part of the reservation.

Results are shared with you. Billing for the term starts when the reservation passes acceptance.

Acceptance sequence. Schematic · not to scale.

Before the term starts, every node and the cluster as a whole pass five tests: burn-in under sustained GPU load, NVIDIA DCGM diagnostics at level 3 on every GPU, an NCCL all-reduce test across the reservation, InfiniBand link checks and, where storage is reserved, a storage throughput check. The results are shared with you. Billing for the term starts when the reservation passes acceptance. Connections: Acceptance tests to Results; Results to Acceptance; Acceptance to Billing starts.

Acceptance testsBefore the term
  • 01Burn-in

    On every node, for theduration in your orderform

  • 02Diagnostics
    $ dcgmi diag -r 3

    On every GPU

  • 03All-reduce
    $ all_reduce_perf

    Across the fullreservation

  • 04Link checks
    speed
    rated
    errors
    clean

    On every InfiniBandport

  • 05Storage
    check
    throughput

    Where storage is partof the reservation

  • RecordResultsShared with you
    burn-in
    every node
    diagnostics
    every GPU
    bus bandwidth
    per node pair
    links
    every port
    storage
    if reserved
  • Acceptance (highlighted)

    The reservation passes

    Every node, and the cluster as awhole

  • Billing startsFor the term
Acceptance testsBefore the term
  • 01Burn-in

    On every node, for theduration in your orderform

  • 02Diagnostics
    $ dcgmi diag -r 3

    On every GPU

  • 03All-reduce
    $ all_reduce_perf

    Across the fullreservation

  • 04Link checks
    speed
    rated
    errors
    clean

    On every InfiniBand port

  • 05Storage
    check
    throughput

    Where storage is part of thereservation

  • RecordResultsShared with you
    burn-in
    every node
    diagnostics
    every GPU
    bus bandwidth
    per node pair
    links
    every port
    storage
    if reserved
  • Acceptance (highlighted)

    The reservation passes

    Every node, and the cluster as a whole

  • Billing startsFor the term
  • 01Burn-in

    On every node, for the duration in yourorder form

  • 02Diagnostics
    $ dcgmi diag -r 3

    On every GPU

  • 03All-reduce
    $ all_reduce_perf

    Across the full reservation

  • 04Link checks
    speed
    rated
    errors
    clean

    On every InfiniBand port

  • 05Storage
    check
    throughput

    Where storage is part of the reservation

  • RecordResultsShared with you
    burn-in
    every node
    diagnostics
    every GPU
    bus bandwidth
    per node pair
    links
    every port
    storage
    if reserved
  • Acceptance (highlighted)

    The reservation passes

    Every node, and the cluster as a whole

  • Billing startsFor the term
Acceptance tests
  • Billing follows acceptance

5. Your verification

After acceptance you have an acceptance window, stated in your order form, to run your own NCCL and training benchmarks. If the cluster does not meet the disclosed specification, Fantasti remediates before the term starts.

6. Node replacement

A node that fails during the term is replaced within the replacement window stated in your order form, counted from the failure being confirmed. Time a node is unavailable is credited at the reservation rate.

7. Support

Response-time targets are set by severity.

TAB 01Response-time targets, by severity Source · Your order form
Response-time targets, by severity
SeverityMeaningResponse-time target
Severity 1Cluster unusableStated in your order form
Severity 2DegradedStated in your order form
Severity 3Question or minor issueStated in your order form

Every reservation has a named escalation contact and a shared channel. The named contact is an AI agent. Every agent identifies itself as an AI in its first message and in its signature. These targets are separate from the uptime commitment in the SLA.

8. Disclosure

The region and facility behind the reservation are named in the order form. Where the capacity layer requires it, disclosure is under NDA.

9. Not covered

Spot capacity, on-demand instances, Sandboxes, and networking between fabrics (InfiniBand does not cross fabrics).

10. Versions

Changes are published as a new version at a new address. Version history: v1, first publication.