Compute Private Preview
One GPU,
eight GPUs,
or a rack of 72.
Virtual machines with one NVIDIA GPU or a whole eight-GPU board, on demand, on spot or reserved, and CPU machines for the services around them. GB300 NVL72 and GB200 NVL72 racks come on reserved terms. All of it is metered by the second, on one statement.
fx.instances.create(gpu="H100:1") - 1. SXM6 module · 8 per baseboard, four per side
- 2. Heatsink · one per module; this one is drawn lifted
- 3. B200 GPU · 180 GB HBM3E, one per module under the heatsink
- 4. NVLink Switch chip · 2 per baseboard, in the board centre
- 5. Baseboard · joins the 8 GPUs in one NVLink domain, NVLink 5 at 1.8 TB/s per GPU
- GPU, on demand
- From $1.86 per GPU-hour
- GPU, on spot
- From $0.87 per GPU-hour
- CPU
- $0.0396 per vCPU-hour
- Reserved and racks
- Request a quote
Compute at a glance
- GPU Instances 1 or 8 GPUs B300, B200, H200, H100, RTX PRO 6000 or L40S in one virtual machine.
- CPU Instances 2 to 256 vCPU General-purpose machines with 4 GiB of memory per vCPU.
- Rack-scale By request 72 GPUs per rack GB300 NVL72 and GB200 NVL72, reserved.
- Metering By the second Billed hourly, on one statement.
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
gpu = fx.instances.create(
name="ft-01",
gpu="H100:1", # or "B200:8" for a whole board
capacity=fantasti.OnDemand(), # or Spot(...), or Reserved("rsv_…")
volumes={"/data": "vol_datasets"},
)
cpu = fx.instances.create(name="prep-01", cpu=32) # 32 vCPU · 128 GiB
gpu.wait("ready")
print(gpu.placement) # region, fabric, rate, capacity mode Output · illustrative
region us-east
fabric b
rate catalog rate · per GPU-hour
capacity on-demand
id ins_8f2a… $ fantasti instance create ft-01 --gpu H100:1 --on-demand
$ fantasti instance create prep-01 --cpu 32
$ fantasti instance list Output · illustrative
NAME SHAPE CAPACITY PLACED STATE
ft-01 H100:1 on-demand us-east · fabric b running
prep-01 32 vCPU on-demand us-east running curl -X POST https://api.fantasti.ai/v1/instances \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "name": "ft-01", "gpu": "H100:1", "capacity": { "mode": "on_demand" } }'
# { "id": "ins_8f2a…", "status": "placing" } Create a machine. Read where it landed.
Ask for a GPU and a capacity mode. The Orchestrator places the instance and returns a record of where it runs and what it costs.
What comes back
- Region and fabricwhere the instance runs
- Ratethe rate per GPU-hour for this placement
- Capacity modeon-demand, spot or reserved
- Idsfor your scheduler and your logs
Two commands create a GPU instance, ft-01 with H100:1, and a CPU instance, prep-01 with 32 vCPU and 128 GiB. The Fantasti Orchestrator considers 4 pools and places ft-01 in us-east · fabric b, on-demand. Both machines run in one private network, the volume vol_datasets is mounted at /data, and both are metered by the second onto one statement. Connections: fantasti to Orchestrator: create; Orchestrator to ft-01: placed; ft-01 to prep-01; prep-01 to Statement: meter.
$ fantasti instance createft-01 --gpu H100:1--on-demandft-01 placingft-01 ready$ fantasti instance createprep-01 --cpu 32prep-01 ready- request
- H100:1 · on-demand
4 pools considered- placed
- us-east · fabric b
- request
- 32 vCPU · on-demand
- placed
- us-east
- meter
- starts when ready
- 1 × H100
- region
- us-east
- fabric
- b
- capacity
- on-demand
- memory
- 80 GB HBM3
- /data
- vol_datasets
- id
- ins_8f2a…
- memory
- 128 GiB
- capacity
- on-demand
- ft-01 · H100:1
- $5.40 per GPU-hour
- prep-01 · 32 vCPU
- $1.2672 per hour
- vol_datasets · SSD
- $0.0888 per GiB-month
Metered by the second, billed hourly.
$ fantasti instance create ft-01 --gpuH100:1 --on-demandft-01 placingft-01 ready$ fantasti instance create prep-01--cpu 32prep-01 ready- request
- H100:1 · on-demand
4 pools considered- placed
- us-east · fabric b
- request
- 32 vCPU · on-demand
- placed
- us-east
- meter
- starts when ready
- 1 × H100
- region
- us-east
- fabric
- b
- capacity
- on-demand
- memory
- 80 GB HBM3
- /data
- vol_datasets
- id
- ins_8f2a…
- memory
- 128 GiB
- capacity
- on-demand
- ft-01 · H100:1
- $5.40 per GPU-hour
- prep-01 · 32 vCPU
- $1.2672 per hour
- vol_datasets · SSD
- $0.0888 per GiB-month
Metered by the second,billed hourly.
- Placed
- Your request
- Metered
- Private network
GPU Instances.
A virtual machine with one GPU, or with all eight GPUs of an HGX board joined by NVLink. Each GPU is passed through whole. Nothing is sliced or shared.
- NVLink 5
B300 270 GB HBM3E $10.45 /GPU·hr per GPU-hour, on demand
- NVLink 5
B200 180 GB HBM3E $9.35 /GPU·hr per GPU-hour, on demand
- NVLink 4
H200 SXM 141 GB HBM3E $5.94 /GPU·hr per GPU-hour, on demand
- NVLink 4
H100 SXM5 80 GB HBM3 $5.40 /GPU·hr per GPU-hour, on demand
- PCIe Gen5
RTX PRO 6000 96 GB GDDR7 $2.16 /GPU·hr per GPU-hour, on demand
- PCIe Gen4
L40S 48 GB GDDR6 $1.86 /GPU·hr per GPU-hour, on demand
| GPU | 1-GPU shape | 8-GPU shape | Between GPUs | Joins a cluster | Status | On-demand |
|---|---|---|---|---|---|---|
| B300 | 24 vCPU · 346 GiB | 192 vCPU · 2,768 GiB | NVLink 5 · 8 GPUs per HGX board | Yes | Private Preview | $10.45 /GPU·hr |
| B200 | 20 vCPU · 224 GiB | 160 vCPU · 1,792 GiB | NVLink 5 · 8 GPUs per HGX board | Yes | Private Preview | $9.35 /GPU·hr |
| H200 | 16 vCPU · 200 GiB | 128 vCPU · 1,600 GiB | NVLink 4 · 8 GPUs per HGX board | Yes | Private Preview | $5.94 /GPU·hr |
| H100 | 16 vCPU · 200 GiB | 128 vCPU · 1,600 GiB | NVLink 4 · 8 GPUs per HGX board | Yes | Private Preview | $5.40 /GPU·hr |
| RTX PRO 6000 | 24 vCPU · 218 GiB | 192 vCPU · 1,744 GiB | No NVLink · PCIe 5.0 x16 | No | Private Preview | $2.16 /GPU·hr |
| L40S | 8 to 40 vCPU · 32 to 160 GiB | 1 to 4 GPUs per machine 16 to 192 vCPU · 96 to 1,152 GiB | No NVLink · PCIe Gen4 x16 | No | Private Preview | $1.86 /GPU·hr |
- USD per GPU-hour · preview list price · applies when your account opens · reviewed 2026-10-10
- L40S. List price for the base shape: 1 GPU, 8 vCPU, 32 GiB. Larger shapes are quoted.
- Local NVMe is included in the node price.
- 01Specifications are NVIDIA reference figures.
- 02RTX PRO 6000 and L40S machines have no InfiniBand adapters. They run as single instances.
- 03L40S machines hold 1 to 4 GPUs.
CPU machines for everything around the GPUs.
Data preparation, tokenizing, evaluation harnesses, schedulers, login nodes and services do not need a GPU. CPU Instances run in the same private network as your GPU nodes and appear on the same bill.
$0.0396 per vCPU-hour
Includes 4 GiB of memory per vCPU. Metered by the second, billed hourly.
2 vCPU$0.0792 /hr to 256 vCPU$10.1376 /hr
| vCPU | Memory, GiB | Availability | Per instance-hour | Per month |
|---|---|---|---|---|
| 2 | 8 | All regions | $0.0792 /hr | $57.82 |
| 4 | 16 | All regions | $0.1584 /hr | $115.63 |
| 8 | 32 | All regions | $0.3168 /hr | $231.26 |
| 16 | 64 | All regions | $0.6336 /hr | $462.53 |
| 32 | 128 | All regions | $1.2672 /hr | $925.06 |
| 48 | 192 | All regions | $1.9008 /hr | $1,387.58 |
| 64 | 256 | All regions | $2.5344 /hr | $1,850.11 |
| 96 | 384 | All regions | $3.8016 /hr | $2,775.17 |
| 128 | 512 | All regions | $5.0688 /hr | $3,700.22 |
| 160 | 640 | Selected regions | $6.3360 /hr | $4,625.28 |
| 192 | 768 | Selected regions | $7.6032 /hr | $5,550.34 |
| 224 | 896 | Selected regions | $8.8704 /hr | $6,475.39 |
| 256 | 1,024 | Selected regions | $10.1376 /hr | $7,400.45 |
- USD, memory included · a month is 730 hours · preview list price · applies when your account opens · reviewed 2026-10-10
- AMD EPYC 9654, x86-64. On-demand capacity.
fx.instances.create(name="prep-01", cpu=32) # 32 vCPU · 128 GiB What teams run on them
- 01Workspace head nodes, which stay on-demand while GPU workers come and go.
- 02A small CPU pool in a GPU cluster, so system services stay off GPU nodes and GPU pools can scale to zero.
- 03Data loaders, ETL and evaluation harnesses.
- 04Gateways, jump hosts and VPN endpoints.
- Sheet
- 01 / 02
- Title
- Instance shapes
- Tables
- TAB 01 · TAB 02
- Reviewed
- 2026-10-10
Rack-scale: 72 GPUs, one NVLink domain.
GB300 NVL72 and GB200 NVL72 racks join 72 GPUs and 36 Grace CPUs in one NVLink domain: 18 compute trays and 9 NVLink switch trays in a liquid-cooled rack.
Rack-scale systems are reserved capacity. Talk to us about term, delivery and configuration.
- GB300 NVL7272 Blackwell Ultra GPUs · 279 GB HBM3E per GPU · 20 TB in the rack
- GB200 NVL7272 Blackwell GPUs · 186 GB HBM3E per GPU · 13.4 TB in the rack
- NVLink 51.8 TB/s per GPU · 130 TB/s across the rack
- 1. Power shelf · 8 per rack
- 2. Management switch · 2 per rack
- 3. Compute tray · 18 × 1RU, 2 Grace CPUs and 4 GPUs each
- 4. NVLink switch tray · 9 × 1RU
- 5. NVLink switch chip · 2 per switch tray
- 6. Grace CPU · 2 per compute tray, 36 per rack
- 7. Blackwell Ultra GPU · 4 per compute tray, 72 per rack
- 8. ConnectX-8 SuperNIC, 800 Gb/s · 4 per compute tray
- 9. Liquid-cooling manifold · rear, supply and return
- 10. Bus bar · rear, fed by the power shelves
On demand, on spot or reserved.
- 01 On-demand From $1.86 per GPU-hourThe published rate per GPU‑hour while the machine runs. No minimum term.
fantasti.OnDemand() - 02 Spot From $0.87 per GPU-hourThe market's spot price, with a max price if you want one. Capacity can be reclaimed. GPU machines, not CPU machines.
fantasti.Spot(max_price=2.40) - 03 Reserved Request a quoteA term commitment for guaranteed capacity. Rack‑scale systems are reserved.
fantasti.Reserved("rsv_7d2k")
The spot price always stays below the on-demand rate for the same GPU. "From" is the lowest spot price for that GPU. The spot price moves with supply and demand.
Instance, cluster or rack?
| Question | GPU Instances | GPU Clusters | Rack-scale |
|---|---|---|---|
| Unit | One machine with 1 or 8 GPUs | Many 8-GPU nodes | NVL72 racks, 72 GPUs each |
| Between GPUs | NVLink inside the machine | NVLink inside a node, InfiniBand between nodes | One NVLink domain across the rack |
| GPUs | B300, B200, H200, H100, RTX PRO 6000, L40S | B300, B200, H200, H100 | GB300 NVL72, GB200 NVL72 |
| Capacity | On-demand, spot, reserved | On-demand, spot, reserved | Reserved |
| Status | Private Preview | Private Preview | By request |
-
GPU Instances
- Unit
- One machine with 1 or 8 GPUs
- Between GPUs
- NVLink inside the machine
- GPUs
- B300, B200, H200, H100, RTX PRO 6000, L40S
- Capacity
- On-demand, spot, reserved
- Status
- Private Preview
-
GPU Clusters
- Unit
- Many 8-GPU nodes
- Between GPUs
- NVLink inside a node, InfiniBand between nodes
- GPUs
- B300, B200, H200, H100
- Capacity
- On-demand, spot, reserved
- Status
- Private Preview
-
Rack-scale
- Unit
- NVL72 racks, 72 GPUs each
- Between GPUs
- One NVLink domain across the rack
- GPUs
- GB300 NVL72, GB200 NVL72
- Capacity
- Reserved
- Status
- By request
- Sheet
- 02 / 02
- Title
- Instance, cluster or rack
- Tables
- TAB 03
- Reviewed
- 2026-10-10
Questions about Compute.
Do I get root?
Can I switch an instance between on-demand and spot?
What happens to my disk when I stop an instance?
Does local NVMe cost extra?
Which processor do CPU Instances run on?
Am I billed while an instance is being placed?
Which regions can I use?
Can a rack run on demand?
Tell us the GPU, the count, the term and the start date.
Companies can request a place in the first cohort.