Compute Private Preview

One GPU,
eight GPUs,
or a rack of 72.

Virtual machines with one NVIDIA GPU or a whole eight-GPU board, on demand, on spot or reserved, and CPU machines for the services around them. GB300 NVL72 and GB200 NVL72 racks come on reserved terms. All of it is metered by the second, on one statement.

fx.instances.create(gpu="H100:1")
Opens with the first cohort.Rack-scale systems are by request.
DWG 01HGX B200 8-GPU baseboard Source · NVIDIA HGX B200 documentation
NVIDIA HGX B200 8-GPU baseboard, dimetric shop drawing
  1. 1. SXM6 module · 8 per baseboard, four per side
  2. 2. Heatsink · one per module; this one is drawn lifted
  3. 3. B200 GPU · 180 GB HBM3E, one per module under the heatsink
  4. 4. NVLink Switch chip · 2 per baseboard, in the board centre
  5. 5. Baseboard · joins the 8 GPUs in one NVLink domain, NVLink 5 at 1.8 TB/s per GPU
GPU, on demand
From $1.86 per GPU-hour
GPU, on spot
From $0.87 per GPU-hour
CPU
$0.0396 per vCPU-hour
Reserved and racks
Request a quote
§02 Create Preview API
Preview API · subject to change
import fantasti

fx = fantasti.Client()                          # reads FANTASTI_API_KEY

gpu = fx.instances.create(
    name="ft-01",
    gpu="H100:1",                               # or "B200:8" for a whole board
    capacity=fantasti.OnDemand(),               # or Spot(...), or Reserved("rsv_…")
    volumes={"/data": "vol_datasets"},
)
cpu = fx.instances.create(name="prep-01", cpu=32)   # 32 vCPU · 128 GiB

gpu.wait("ready")
print(gpu.placement)                            # region, fabric, rate, capacity mode

Output · illustrative

region     us-east
fabric     b
rate       catalog rate · per GPU-hour
capacity   on-demand
id         ins_8f2a…
$ fantasti instance create ft-01 --gpu H100:1 --on-demand
$ fantasti instance create prep-01 --cpu 32
$ fantasti instance list

Output · illustrative

NAME      SHAPE      CAPACITY    PLACED                STATE
ft-01     H100:1     on-demand   us-east · fabric b    running
prep-01   32 vCPU    on-demand   us-east               running
curl -X POST https://api.fantasti.ai/v1/instances \
  -H "Authorization: Bearer $FANTASTI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "name": "ft-01", "gpu": "H100:1", "capacity": { "mode": "on_demand" } }'
# { "id": "ins_8f2a…", "status": "placing" }

Create a machine. Read where it landed.

Ask for a GPU and a capacity mode. The Orchestrator places the instance and returns a record of where it runs and what it costs.

What comes back

  • Region and fabricwhere the instance runs
  • Ratethe rate per GPU-hour for this placement
  • Capacity modeon-demand, spot or reserved
  • Idsfor your scheduler and your logs

How capacity is placed

What the sample creates. Illustrative.

Two commands create a GPU instance, ft-01 with H100:1, and a CPU instance, prep-01 with 32 vCPU and 128 GiB. The Fantasti Orchestrator considers 4 pools and places ft-01 in us-east · fabric b, on-demand. Both machines run in one private network, the volume vol_datasets is mounted at /data, and both are metered by the second onto one statement. Connections: fantasti to Orchestrator: create; Orchestrator to ft-01: placed; ft-01 to prep-01; prep-01 to Statement: meter.

Private networkus-east
  • 01Terminalfantasti
    $ fantasti instance create ft-01 --gpu H100:1 --on-demandft-01 placingft-01 ready$ fantasti instance create prep-01 --cpu 32prep-01 ready
  • 02OrchestratorPlaces
    request
    H100:1 · on-demand
    4 pools considered
    placed
    us-east · fabric b
    request
    32 vCPU · on-demand
    placed
    us-east
    meter
    starts when ready
  • ft-01H100:1
    1 × H100
    region
    us-east
    fabric
    b
    capacity
    on-demand
    memory
    80 GB HBM3
    /data
    vol_datasets
    id
    ins_8f2a…
  • prep-0132 vCPU
    memory
    128 GiB
    capacity
    on-demand
  • 03StatementPer second
    ft-01 · H100:1
    $5.40 per GPU-hour
    prep-01 · 32 vCPU
    $1.2672 per hour
    vol_datasets · SSD
    $0.0888 per GiB-month

    Metered by the second, billed hourly.

  • 01Terminalfantasti
    $ fantasti instance create ft-01 --gpu H100:1 --on-demandft-01 placingft-01 ready$ fantasti instance create prep-01 --cpu 32prep-01 ready
  • 02OrchestratorPlaces
    request
    H100:1 · on-demand
    4 pools considered
    placed
    us-east · fabric b
    request
    32 vCPU · on-demand
    placed
    us-east
    meter
    starts when ready
  • ft-01H100:1
    1 × H100
    region
    us-east
    fabric
    b
    capacity
    on-demand
    memory
    80 GB HBM3
    /data
    vol_datasets
    id
    ins_8f2a…
  • prep-0132 vCPU
    memory
    128 GiB
    capacity
    on-demand
  • 03StatementPer second
    ft-01 · H100:1
    $5.40 per GPU-hour
    prep-01 · 32 vCPU
    $1.2672 per hour
    vol_datasets · SSD
    $0.0888 per GiB-month

    Metered by the second,billed hourly.

Private network
  • Placed
  • Your request
  • Metered
  • Private network
§03 GPU Instances Shapes reviewed 2026-10-10

GPU Instances.

A virtual machine with one GPU, or with all eight GPUs of an HGX board joined by NVLink. Each GPU is passed through whole. Nothing is sliced or shared.

TAB 01GPU instance shapes Source · Fantasti platform configuration Reviewed 2026-10-10
GPU instance shapes. Virtual machine shapes per GPU, and the preview list price in USD per GPU-hour, reviewed 2026-10-10.
GPU 1-GPU shape 8-GPU shape Between GPUs Joins a cluster Status On-demand
B300 24 vCPU · 346 GiB 192 vCPU · 2,768 GiB NVLink 5 · 8 GPUs per HGX board Yes Private Preview $10.45 /GPU·hr
B200 20 vCPU · 224 GiB 160 vCPU · 1,792 GiB NVLink 5 · 8 GPUs per HGX board Yes Private Preview $9.35 /GPU·hr
H200 16 vCPU · 200 GiB 128 vCPU · 1,600 GiB NVLink 4 · 8 GPUs per HGX board Yes Private Preview $5.94 /GPU·hr
H100 16 vCPU · 200 GiB 128 vCPU · 1,600 GiB NVLink 4 · 8 GPUs per HGX board Yes Private Preview $5.40 /GPU·hr
RTX PRO 6000 24 vCPU · 218 GiB 192 vCPU · 1,744 GiB No NVLink · PCIe 5.0 x16 No Private Preview $2.16 /GPU·hr
L40S 8 to 40 vCPU · 32 to 160 GiB 1 to 4 GPUs per machine 16 to 192 vCPU · 96 to 1,152 GiB No NVLink · PCIe Gen4 x16 No Private Preview $1.86 /GPU·hr
  • USD per GPU-hour · preview list price · applies when your account opens · reviewed 2026-10-10
  • L40S. List price for the base shape: 1 GPU, 8 vCPU, 32 GiB. Larger shapes are quoted.
  • Local NVMe is included in the node price.
  • 01Specifications are NVIDIA reference figures.
  • 02RTX PRO 6000 and L40S machines have no InfiniBand adapters. They run as single instances.
  • 03L40S machines hold 1 to 4 GPUs.
§04 CPU Instances Private Preview

CPU machines for everything around the GPUs.

Data preparation, tokenizing, evaluation harnesses, schedulers, login nodes and services do not need a GPU. CPU Instances run in the same private network as your GPU nodes and appear on the same bill.

$0.0396 per vCPU-hour

Includes 4 GiB of memory per vCPU. Metered by the second, billed hourly.

2 vCPU$0.0792 /hr to 256 vCPU$10.1376 /hr

TAB 02CPU instance shapes and prices Source · Fantasti price list Reviewed 2026-10-10
CPU instance shapes and prices. 4 GiB per vCPU on every shape. Processor: AMD EPYC 9654, x86-64. Preview list price in USD per instance-hour and per month of 730 hours, reviewed 2026-10-10.
vCPU Memory, GiB Availability Per instance-hour Per month
2 8 All regions $0.0792 /hr $57.82
4 16 All regions $0.1584 /hr $115.63
8 32 All regions $0.3168 /hr $231.26
16 64 All regions $0.6336 /hr $462.53
32 128 All regions $1.2672 /hr $925.06
48 192 All regions $1.9008 /hr $1,387.58
64 256 All regions $2.5344 /hr $1,850.11
96 384 All regions $3.8016 /hr $2,775.17
128 512 All regions $5.0688 /hr $3,700.22
160 640 Selected regions $6.3360 /hr $4,625.28
192 768 Selected regions $7.6032 /hr $5,550.34
224 896 Selected regions $8.8704 /hr $6,475.39
256 1,024 Selected regions $10.1376 /hr $7,400.45
  • USD, memory included · a month is 730 hours · preview list price · applies when your account opens · reviewed 2026-10-10
  • AMD EPYC 9654, x86-64. On-demand capacity.
CPU prices on the price list
fx.instances.create(name="prep-01", cpu=32) # 32 vCPU · 128 GiB

What teams run on them

  • 01Workspace head nodes, which stay on-demand while GPU workers come and go.
  • 02A small CPU pool in a GPU cluster, so system services stay off GPU nodes and GPU pools can scale to zero.
  • 03Data loaders, ETL and evaluation harnesses.
  • 04Gateways, jump hosts and VPN endpoints.
Sheet
01 / 02
Title
Instance shapes
Tables
TAB 01 · TAB 02
Reviewed
2026-10-10
FILM 02Liquid-cooling manifold Illustrative
§05 Rack-scale Reserved terms
By request

Rack-scale: 72 GPUs, one NVLink domain.

GB300 NVL72 and GB200 NVL72 racks join 72 GPUs and 36 Grace CPUs in one NVLink domain: 18 compute trays and 9 NVLink switch trays in a liquid-cooled rack.

Rack-scale systems are reserved capacity. Talk to us about term, delivery and configuration.

  • GB300 NVL7272 Blackwell Ultra GPUs · 279 GB HBM3E per GPU · 20 TB in the rack
  • GB200 NVL7272 Blackwell GPUs · 186 GB HBM3E per GPU · 13.4 TB in the rack
  • NVLink 51.8 TB/s per GPU · 130 TB/s across the rack

Talk to us about a rack

DWG 03GB300 NVL72 rack Source · NVIDIA DGX GB rack documentation
NVIDIA GB300 NVL72 rack, dimetric shop drawing
  1. 1. Power shelf · 8 per rack
  2. 2. Management switch · 2 per rack
  3. 3. Compute tray · 18 × 1RU, 2 Grace CPUs and 4 GPUs each
  4. 4. NVLink switch tray · 9 × 1RU
  5. 5. NVLink switch chip · 2 per switch tray
  6. 6. Grace CPU · 2 per compute tray, 36 per rack
  7. 7. Blackwell Ultra GPU · 4 per compute tray, 72 per rack
  8. 8. ConnectX-8 SuperNIC, 800 Gb/s · 4 per compute tray
  9. 9. Liquid-cooling manifold · rear, supply and return
  10. 10. Bus bar · rear, fed by the power shelves
§07 Which shape Reviewed 2026-10-10

Instance, cluster or rack?

TAB 03Instance, cluster or rack Source · Fantasti Reviewed 2026-10-10
Instance, cluster or rack
Question GPU Instances GPU Clusters Rack-scale
Unit One machine with 1 or 8 GPUsMany 8-GPU nodesNVL72 racks, 72 GPUs each
Between GPUs NVLink inside the machineNVLink inside a node, InfiniBand between nodesOne NVLink domain across the rack
GPUs B300, B200, H200, H100, RTX PRO 6000, L40SB300, B200, H200, H100GB300 NVL72, GB200 NVL72
Capacity On-demand, spot, reservedOn-demand, spot, reservedReserved
Status Private PreviewPrivate PreviewBy request
  • GPU Instances

    Unit
    One machine with 1 or 8 GPUs
    Between GPUs
    NVLink inside the machine
    GPUs
    B300, B200, H200, H100, RTX PRO 6000, L40S
    Capacity
    On-demand, spot, reserved
    Status
    Private Preview
  • GPU Clusters

    Unit
    Many 8-GPU nodes
    Between GPUs
    NVLink inside a node, InfiniBand between nodes
    GPUs
    B300, B200, H200, H100
    Capacity
    On-demand, spot, reserved
    Status
    Private Preview
  • Rack-scale

    Unit
    NVL72 racks, 72 GPUs each
    Between GPUs
    One NVLink domain across the rack
    GPUs
    GB300 NVL72, GB200 NVL72
    Capacity
    Reserved
    Status
    By request
Sheet
02 / 02
Title
Instance, cluster or rack
Tables
TAB 03
Reviewed
2026-10-10
§08 Questions

Questions about Compute.

Do I get root?

Yes. You have root inside your virtual machine.

Are the GPUs shared?

No. Each GPU is passed through whole to one machine.

Can I switch an instance between on-demand and spot?

No. Capacity mode is set when an instance is created. Create a new instance and attach the same volumes.

What happens to my disk when I stop an instance?

Volumes are kept and billed as storage. Local NVMe is erased.

Does local NVMe cost extra?

No. Local NVMe is included in the node price.

Which processor do CPU Instances run on?

AMD EPYC 9654, x86-64. Every shape has 4 GiB of memory per vCPU.

Am I billed while an instance is being placed?

No. Billing starts when the instance is ready.

Which regions can I use?

Regions are confirmed with your quote.

Can a rack run on demand?

No. Rack-scale systems are reserved capacity.
§09 Request Private Preview