ComputeGPU Clusters Private Preview
Many nodes,
one InfiniBand
fabric.
GPU Clusters are groups of 8-GPU nodes joined by InfiniBand and placed together on a single fabric. Use them as plain machines over SSH, or as an isolated cluster with its own Kubernetes API. Pay on demand, on spot or on reserved terms.
fx.clusters.create(gpu="B200:8", nodes=16) One InfiniBand fabric with three switch tiers: root switches, core switches and leaf switches. Your cluster of 16 nodes, 8 GPUs each, sits inside its own partition under neighbouring leaf switches. One node is drawn in full: 8 B200 GPUs on an NVLink baseboard, each GPU with its own InfiniBand port at 400 Gb/s, eight links to its leaf switch. A second cluster on the same fabric sits inside another partition, and no InfiniBand path joins the two. Connections: Core switch to Root switch; Leaf switch to Core switch; Nodes 02–16 to Leaf switch, 8 links; Their nodes to Leaf switch, 8 links; gpu-node-01 to Leaf switch, 8 links; gpu-node-01 to Another cluster.
fx.clusters.create(gpu="B200:8", nodes=16)- gpus
- 8 × B200 · NVLink
- ib0–7
- 8 × 400 Gb/s
- host
- 160 vCPU · 1,792 GiB
- each
- 8 GPUs · 8 ports
No InfiniBand path toyour partition.
fx.clusters.create(gpu="B200:8", nodes=16)- gpus
- 8 × B200 · NVLink
- ib0–7
- 8 × 400 Gb/s
- host
- 160 vCPU · 1,792 GiB
- each
- 8 GPUs · 8 ports
- gpus
- 8 × B200 · NVLink
- ib0–7
- 8 × 400 Gb/s
- host
- 160 vCPU · 1,792 GiB
- each
- 8 GPUs · 8 ports
- One node, eight InfiniBand links
- Not every link is drawn
- No path between partitions
GPU Clusters at a glance
- Node 8 GPUs H100, H200, B200 or B300 on an NVLink baseboard.
- Fabric 400 or 800 Gb/s per GPU InfiniBand with GPUDirect RDMA. 800 Gb/s on B300.
- Placement One fabric per cluster Every node of a cluster sits on the same fabric.
- Isolation A partition per cluster Your InfiniBand traffic is separated from every other cluster.
- Interface Kubernetes or SSH An isolated cluster with its own API, or plain machines.
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
cluster = fx.clusters.create(
name="pretrain-a",
gpu="B200:8", # 8-GPU nodes
nodes=16, # all on one InfiniBand fabric
capacity=fantasti.Reserved("rsv_7d2k"), # or OnDemand(), or Spot(...)
interfaces=["kubernetes"],
mounts={"/data": "fs_datasets"}, # shared filesystem on every node
)
cluster.wait("ready")
cluster.kubeconfig.save("~/.kube/fantasti-acme.yaml") curl -X POST https://api.fantasti.ai/v1/clusters \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "acme-train",
"gpu": "H200:8",
"nodes": 2,
"capacity": { "mode": "spot", "max_price": "3.10", "currency": "USD", "unit": "gpu_hour" },
"interfaces": ["kubernetes", "skypilot"]
}' Response · illustrative
{ "id": "cl_8f2k…", "status": "placing", "fabric": "single" } # Console: Clusters → acme → Connect → download kubeconfig
export KUBECONFIG="$HOME/.kube/fantasti-acme.yaml"
kubectl get pods -A # your workloads and your cluster's system pods, nothing else
kubectl get nodes -L nvidia.com/gpu.product,nvidia.com/gpu.count Output · illustrative
NAME STATUS ROLES GPU.PRODUCT GPU.COUNT
gpu-node-01 Ready <none> NVIDIA-H200 8
gpu-node-02 Ready <none> NVIDIA-H200 8 {
"mcpServers": {
"fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
}
} Clusters tool group · planned
clusters_list
clusters_get
clusters_create asks for confirmation
clusters_delete asks for confirmation gpu- H100, H200, B200 or B300, eight to a node.
nodes- Placed together, on one fabric.
capacity- On-demand, spot or a reservation.
interfaces- Kubernetes, or plain machines over SSH.
A cluster is one request.
Ask for a GPU type and a node count. The Orchestrator looks for a fabric with room for all of them and places the cluster there, or tells you it cannot. A cluster is never split across fabrics.
Planned Cluster tools for your agents, through the MCP server.
Node shapes for clusters.
| GPU | Node shape | Host CPU | InfiniBand per GPU | InfiniBand per node | Host network adapter | Local NVMe |
|---|---|---|---|---|---|---|
| B300 | 192 vCPU · 2,768 GiB | Intel Xeon 6776P | 800 Gb/s | Not published | 400 Gb/s | 6 × 3.84 TBIncluded |
| B200 | 160 vCPU · 1,792 GiB | Intel Xeon Platinum 8580 or 8570 | 400 Gb/s | 3.2 Tb/s | 400 Gb/s | None |
| H200 | 128 vCPU · 1,600 GiB | Intel Xeon Platinum 8468 | 400 Gb/s | 3.2 Tb/s | 200 Gb/s | None |
| H100 | 128 vCPU · 1,600 GiB | Intel Xeon Platinum 8468 | 400 Gb/s | 3.2 Tb/s | 100 Gb/s | None |
- 01Host network adapter is the adapter's rating, not a throughput guarantee.
- 02Specifications are NVIDIA reference figures and node configurations reviewed 2026-10-10.
- 03RTX PRO 6000 and L40S machines have no InfiniBand adapters and run as single instances.
- 04Local NVMe is included in the node price.
- Sheet
- 01 / 02
- Title
- Node shapes for clusters
- Reviewed
- 2026-10-10
InfiniBand per node
3.2 Tb/s InfiniBand per 8-GPU node.
InfiniBand per GPU
- H100,H200,B200
- 400 Gb/s
- B300
- 800 Gb/s
One InfiniBand port for each GPU, 8 to a node.
- 1. GPU under its heatsink · 8 per node, on one baseboard
- 2. InfiniBand adapter · 8 per node, single port
- 3. 400 Gb/s port · 8 × 400 Gb/s = 3.2 Tb/s per node, where the pool provides it
- 4. Chassis · cover and near walls removed
One job, one fabric.
InfiniBand carries traffic between GPUs on different nodes straight from GPU memory to the network adapter, without passing through the host CPU.
A fabric has three switch tiers. Nodes that share a lower tier are closer together, and traffic between them takes fewer hops.
Clusters on Kubernetes expose each node's position in the fabric to your scheduler, so it can keep the ranks of one job close together.
Same core switch: node-1 to node-3 goes up through the leaf tier to the core tier, and down again.
One InfiniBand fabric with three switch tiers and four nodes of one cluster. node-1 and node-2 share a leaf switch. node-3 sits under another leaf switch of the same core switch. node-4 sits under another core switch, reached through the root switch. Nodes that share a lower tier are closer together. Connections: Core switch to Root switch; Leaf switch to Core switch; node-1 to Leaf switch; node-2 to Leaf switch; node-3 to Leaf switch; node-4 to Leaf switch.
- The traced route
- Other links
Isolated at every layer.
What separates your cluster from every other account's.
- 01
Nodes
Cluster nodes are dedicated to you. No node is shared with another account.
- 02
Devices
GPUs and InfiniBand adapters are passed through to your nodes. They are not sliced or shared.
- 03
Fabric
Each cluster has its own InfiniBand partition. Nodes in different clusters cannot reach each other over InfiniBand, even on the same physical fabric.
- 04
Control plane
On Kubernetes, your cluster has its own API server and its own network.
Run your scheduler.
- 01 Early access
Kubernetes
Your own API server and kubeconfig.
kubectl get nodes - 02 Early access
Ray
KubeRay worker groups that scale from zero.
minReplicas: 0 - 03 Early access
SkyPilot
Point your YAML at your cluster's kube context.
infra: k8s/fantasti-acme - 04 Private Preview
SSH
Plain machines, your own stack.
ssh gpu-node-01
Checked before you get it. Watched while you use it.
Before handover, link-state and collective-communication tests run across the cluster's InfiniBand ports, and the results are shared with you.
On Kubernetes clusters, GPU, NVLink and InfiniBand health checks run on every node. A node that fails is cordoned, drained and brought back on healthy hardware.
GPU, InfiniBand and node metrics are collected from the start.
Before handover, link-state and collective-communication tests run across the 128 InfiniBand ports of a 16-node cluster, and the results are shared. While the cluster runs, GPU, NVLink and InfiniBand checks run on every node. In the example one node fails a check: it is cordoned, drained and brought back on healthy hardware. Connections: Handover report to On every node: Handover; On every node to gpu-node-11: A check fails.
- Link state
- 128 of 128 ports up
- Collective communication
- 16 nodes, one job
- Results
- Shared with you
- GPU
- NVLink
- InfiniBand
- pretrain-a
- 16 nodes · 128 GPUs
check failed NVLinkcordoned no new podsdrained pods movedback healthy hardware
- Link state
- 128 of 128 ports up
- Collective communication
- 16 nodes, one job
- Results
- Shared with you
- GPU
- NVLink
- InfiniBand
- pretrain-a
- 16 nodes · 128 GPUs
check failed NVLinkcordoned no new podsdrained pods movedback healthy hardware
- The node that failed a check
Storage every node can see.
| Storage | What it is for | Preview list price |
|---|---|---|
| Shared filesystem$0.1000 per GiB-month | Datasets, code and checkpoints mounted on every node at once. | $0.1000 per GiB-month |
| Object storage$0.0184 per GiB-month · Standard class | S3-compatible buckets for datasets, checkpoints and results. | $0.0184 per GiB-monthStandard class |
| Local NVMeIncluded | 6 × 3.84 TB of scratch space on B300 8-GPU nodes. Erased when the node stops.Not encrypted. | Included |
| Boot volume$0.0888 per GiB-month | A redundant SSD volume per node. | $0.0888 per GiB-month |
On demand, on spot or reserved.
- 01
On-demand
Pay per GPU-hour while the cluster runs. Size depends on what a fabric has free when you ask.
fantasti.OnDemand() - 02
Spot
Run worker nodes at the spot price, with a max price if you want one. If the price rises above your max, those nodes stop and the rest of the cluster keeps running.
fantasti.Spot(max_price=3.10) - 03
Reserved
Commit to a term for guaranteed capacity. A reservation is tied to a GPU type, a region and a fabric, so a reserved cluster always fits on one fabric.
fantasti.Reserved("rsv_7d2k")
| GPU | Memory | Interconnect | Status | On-demand | Reserved | Spot | Action |
|---|---|---|---|---|---|---|---|
| B300 HGX 8-GPU | 270 GB HBM3E | NVLink 5 | Private Preview | $10.45 /GPU·hr | Request a quote | From $1.09 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| B200 HGX 8-GPU | 180 GB HBM3E | NVLink 5 | Private Preview | $9.35 /GPU·hr | Request a quote | From $1.09 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| H200 SXM HGX 8-GPU | 141 GB HBM3E | NVLink 4 | Private Preview | $5.94 /GPU·hr | Request a quote | From $0.87 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| H100 SXM5 HGX 8-GPU | 80 GB HBM3 | NVLink 4 | Private Preview | $5.40 /GPU·hr | Request a quote | From $0.87 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
USD per GPU-hour · preview list price · applies when your account opens · reviewed 2026-10-10
"From" is the lowest spot price for that GPU. The spot price moves with supply and demand. The spot price always stays below the on-demand rate for the same GPU.
- Ways to pay
- Prepaid credit, pay-as-you-go and Enterprise. Each one covers on-demand, spot and reserved capacity.
- Three ways to pay
Instance, cluster or rack?
| Question | GPU Instances | GPU ClustersThis page | Rack-scale |
|---|---|---|---|
| Unit | One machine with 1 or 8 GPUs | Many 8-GPU nodes | NVL72 racks, 72 GPUs each |
| Between GPUs | NVLink inside the machine | NVLink inside a node, InfiniBand between nodes | One NVLink domain across the rack |
| GPUs | B300, B200, H200, H100, RTX PRO 6000, L40S | B300, B200, H200, H100 | GB300 NVL72, GB200 NVL72 |
| Capacity | On-demand, spot, reserved | On-demand, spot, reserved | Reserved |
| Status | Private Preview | Private Preview | By request |
- Sheet
- 02 / 02
- Title
- Storage, terms and shapes
- Unit
- USD per GPU-hour
- Reviewed
- 2026-10-10