§00Fantasti Cloud Platform Private Preview
From one GPU to
the whole rack.
NVIDIA GPUs, CPU machines, clusters and storage behind one API and one bill. The Fantasti Orchestrator is built to place every workload on one fabric and meter it by the second, with AI agents operating it.
Programs
-
NVIDIA Inception program member
- Member of Cloudflare for Startups
Ask for GPUs in plain words.
Tell your coding agent what you need and the most you will pay. It makes one call. The Orchestrator checks your policy, places the work on one fabric and keeps a record of every pool it turned away.
You ask
Plain wordsI need two 8-GPU H200 nodes on one InfiniBand fabric for a fine-tuning run, in a US region. Spot is fine. Do not pay more than $3.10 per GPU-hour.
- gpu
- H200:8
- nodes
- 2
- fabric
- infiniband
- regions
- us-*
- capacity
- spot
- max_price
- 3.10
Your agent runs
One call$ fantasti cluster create acme-train \
--gpu H200:8 --nodes 2 \
--spot --max-price 3.10 The same request in Python, REST and YAML. See the code
The Orchestrator answers
Recordcl_8f2k - Request
- 16 × H200 SXM · InfiniBand · us-* · spot ≤ 3.10
- Policy
- placement: { fabric: infiniband, regions: ["us-*"] }Within policy
| Pool4 checked | Region and fabric | Free GPUsNeeds 16 | Result |
|---|---|---|---|
| pool-b | us-east · fabric b | 8 free on one fabric | Rejected: 8 free < 16 |
| pool-c | us-west · no InfiniBand | Not counted: no InfiniBand fabric | Rejected: no InfiniBand |
| pool-d | eu-west · fabric d | 96 free on one fabric | Rejected: region not allowed |
| pool-a | us-east · fabric a | 64 free on one fabric | Placed |
Where it runs
PlacedPlaced on
us-east · fabric a
16 × H200 SXM, on one InfiniBand fabric
Pricespot · at or below your max 3.10
Recordplaced · us-east · fabric a · 16×H200 SXM · one fabric
I need two 8-GPU H200 nodes on one InfiniBand fabric for a fine-tuning run, in a US region. Spot is fine. Do not pay more than $3.10 per GPU-hour. The request reads as gpu H200:8, nodes 2, fabric infiniband, regions us-*, capacity spot, max_price 3.10. The request was 16 × H200 SXM · InfiniBand · us-* · spot ≤ 3.10. 4 pools were checked. pool-b in us-east, fabric b was rejected: 8 free < 16. pool-c in us-west, no InfiniBand was rejected: no InfiniBand. pool-d in eu-west, fabric d was rejected: region not allowed. The job was placed in us-east, fabric a: 16 H200 SXM GPUs on one InfiniBand fabric. Price: spot ≤ 3.10.
- How the agents work
- MCP server Early access
- How capacity is placed
Everything between the GPU and the invoice.
The Fantasti Cloud Platform is every product on this list, on one account, one set of policies and one statement. Each line is a product, its status and the call that starts it.
Status
- Private Preview
- Early access
- By request
- Planned
01Compute
GPUs and CPUs as instances, clusters and racks.
- GPU Instances Private Preview Virtual machines with one or eight GPUs.
fx.instances.create(gpu="H100:1") - CPU Instances Private Preview General-purpose machines beside your GPUs.
fx.instances.create(cpu=32) - GPU Clusters Private Preview 8-GPU nodes on one InfiniBand fabric.
fx.clusters.create(gpu="B200:8", nodes=16) - Rack-scale By request GB300 NVL72 and GB200 NVL72 racks, reserved.
capacity=Reserved("rsv_7d2k") - Spot Private Preview Market-priced GPUs with a max price you set.
capacity=Spot(max_price=3.10)
02Develop
Environments people and agents work in.
- Workspaces Early access GPU dev environments that scale to Ray clusters.
fx.workspaces.create(gpu="H100:1") - Sandboxes Early access Short-lived isolated environments for AI agents.
fx.sandboxes.create(ttl="15m") - Flows Planned Python workflows from experiment to production, on GPUs.
@fantasti(gpu="H200:8", capacity="spot")
03Data and network
Storage and networking around your compute.
- Storage Private Preview Volumes, shared filesystems and S3-compatible buckets.
fx.storage.filesystems.create(size="20TiB") - Networking Private Preview Private networks, firewalls, load balancers and InfiniBand.
net.security_groups.create(name="train-nodes")
04Run
Ways to run work without managing machines.
- Batch Inference Private Preview Offline inference at market prices, with checkpoints.
fx.batch.create(input="s3://…") - Serverless Early access Run a container as a job or endpoint.
fx.jobs.create(image=…, gpu="H200:8") - Integrations Early access Kubernetes, Ray, SkyPilot and NVIDIA OSMO on your cluster.
infra: k8s/fantasti-acme
05Operate
The control plane and the shared tools.
- Orchestrator Private Preview Policy, placement, pricing and recovery for every workload.
placement: { fabric: infiniband } - Observability Early access GPU metrics, logs and traces, with dashboards.
fx.metrics.query("…")
Access
One console, one API, one bill.
- API Private Preview Python, REST and a CLI over one client.
fx = fantasti.Client() - MCP server Early access The Fantasti API as tools for AI agents.
"mcpServers": { "fantasti": … } - Console and one bill Private Preview One console and one statement for every product.
statement · cost centers
One client for all of it.
Python, a CLI, REST, YAML, or an MCP server your agent calls directly. The same capacity object works on every product: on-demand, reserved, or spot with a max price.
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
cluster = fx.clusters.create(
name="acme-train",
gpu="H200:8", # 8-GPU nodes
nodes=2, # placed on one InfiniBand fabric
capacity=fantasti.Spot(max_price=3.10), # your ceiling, USD per GPU-hour
interfaces=["kubernetes", "skypilot"],
)
cluster.wait("ready")
cluster.kubeconfig.save("~/.kube/fantasti-acme.yaml") # create the cluster
$ fantasti cluster create acme-train --gpu H200:8 --nodes 2 --spot --max-price 3.10
# read why it was placed where it was
$ fantasti explain cl_8f2k {
"mcpServers": {
"fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
}
} Tools · early access
catalog gpus_list
sandboxes sandboxes_create sandboxes_exec sandboxes_read_file
sandboxes_write_file sandboxes_delete
workspaces workspaces_list workspaces_start workspaces_stop
jobs jobs_create jobs_logs
usage usage_get curl -X POST https://api.fantasti.ai/v1/clusters \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "acme-train",
"gpu": "H200:8",
"nodes": 2,
"capacity": { "mode": "spot", "max_price": "3.10", "currency": "USD", "unit": "gpu_hour" },
"interfaces": ["kubernetes", "skypilot"]
}'
# { "id": "cl_8f2k…", "status": "placing", "fabric": "single" } # policy.yaml · applies to every request from team "research"
team: research
placement:
gpus: [H100, H200, B200]
fabric: infiniband # multi-node jobs stay on one fabric
regions: ["us-*"]
capacity:
allow: [on_demand, spot, reserved]
spot:
max_price: { H100: 2.40, H200: 3.10 } # USD per GPU-hour, your ceilings
limits:
gpus: { H200: 32 }
sandboxes: { concurrent: 200, max_ttl: 60m }
spend: { monthly_usd: 40000 }
cost_center: research-llm Output · illustrativefantasti cluster create
- request
- H200:8 × 2 · one InfiniBand fabric · spot ≤ 3.10
- pools
- 4 considered
- placed
- cl_8f2k… · us-east · fabric a
- state
- running · kubeconfig saved
One client, every product
-
fx.instances.create(gpu="H100:1")GPU Instances -
fx.clusters.create(gpu="H200:8", nodes=2)GPU Clusters -
fx.sandboxes.create(ttl="15m")Sandboxes -
fx.batch.create(input="s3://…")Batch Inference -
fx.jobs.create(image=…, gpu="H200:8")Serverless
- API reference
- MCP server Early access
- Capacity object
- One argument decides how the work is paid for, on every product.
- On-demand
fantasti.OnDemand()Work that must not stop and has no fixed term.- Spot
fantasti.Spot(max_price=3.10)Work that checkpoints and can wait.- Reserved
fantasti.Reserved("rsv_7d2k")Capacity you need on a date, guaranteed.
Blackwell, in three sizes.
B300 and B200 on 8-GPU HGX boards with NVLink 5, for training and large-model inference. RTX PRO 6000 Blackwell Server Edition on a single card, for inference, simulation and rendering. Hopper and L40S are in the catalog below.
-
HGX B300 8-GPU baseboard. Source · NVIDIA HGX reference architecture NVIDIA HGX B300 baseboard: 8 SXM modules under heatsinks, four per side, with 2 NVLink Switch chips in the board centre and 8 ConnectX-8 SuperNICs along one edge. One heatsink is lifted to show the GPU package.
Memory per GPU
270 GB
NVIDIA reference figure
- Memory
- 270 GB HBM3E per GPU
- Between GPUs
- NVLink 5
- Board
- 8 GPUs per HGX B300 board
$10.45 /GPU·hrOn-demand, preview list price
Datasheet -
HGX B200 8-GPU baseboard. Source · NVIDIA HGX B200 documentation NVIDIA HGX B200 baseboard: 8 SXM6 modules under heatsinks, four per side, with 2 NVLink Switch chips in the board centre. One heatsink is lifted to show the GPU package.
Memory per GPU
180 GB
NVIDIA reference figure
- Memory
- 180 GB HBM3E per GPU
- Between GPUs
- NVLink 5
- Board
- 8 GPUs per HGX B200 board
$9.35 /GPU·hrOn-demand, preview list price
Datasheet -
RTX PRO 6000 Blackwell Server Edition. Source · NVIDIA datasheet NVIDIA RTX PRO 6000 Blackwell Server Edition: a dual-slot, full-height, full-length PCIe card with a passive heatsink and a PCIe x16 edge connector. The cover is cut away to show the fins.
Memory per GPU
96 GB
NVIDIA reference figure
- Memory
- 96 GB GDDR7 with ECC
- Host link
- PCIe Gen5
- Form
- One card, no NVLink
$2.16 /GPU·hrOn-demand, preview list price
Datasheet
Eight systems, one price list.
Preview list prices in USD per GPU-hour, metered by the second. Rack-scale systems are reserved and quoted.
| GPU | Memory | Interconnect | Status | On-demand | Reserved | Spot | Action |
|---|---|---|---|---|---|---|---|
| B300 HGX 8-GPU | 270 GB HBM3E | NVLink 5 | Private Preview | $10.45 /GPU·hr | Request a quote | From $1.09 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| B200 HGX 8-GPU | 180 GB HBM3E | NVLink 5 | Private Preview | $9.35 /GPU·hr | Request a quote | From $1.09 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| H200 SXM HGX 8-GPU | 141 GB HBM3E | NVLink 4 | Private Preview | $5.94 /GPU·hr | Request a quote | From $0.87 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| H100 SXM5 HGX 8-GPU | 80 GB HBM3 | NVLink 4 | Private Preview | $5.40 /GPU·hr | Request a quote | From $0.87 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| RTX PRO 6000 PCIe · Server Edition | 96 GB GDDR7 | PCIe Gen5 | Private Preview | $2.16 /GPU·hr | Request a quote | From $0.87 Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max.How spot works | Request |
| L40S PCIe | 48 GB GDDR6 | PCIe Gen4 | Private Preview | $1.86 /GPU·hr | Request a quote | Request | |
| GB300 NVL72 Rack-scale · 72 GPUs | 279 GB HBM3E | NVLink 5 · 72 | By request | Reserved · contact sales | Request a quote | Request | |
| GB200 NVL72 Rack-scale · 72 GPUs | 186 GB HBM3E | NVLink 5 · 72 | By request | Reserved · contact sales | Request a quote | Request |
USD per GPU-hour · preview list price · applies when your account opens · reviewed 2026-10-10
"From" is the lowest spot price for that GPU. The spot price moves with supply and demand. The spot price always stays below the on-demand rate for the same GPU.
- Sheet
- 01 / 02
- Title
- Catalog
- Unit
- USD per GPU-hour
- Reviewed
- 2026-10-10
Rack-scale: 72 GPUs in one NVLink domain
72 GPUs in one NVLink domain.
GB300 NVL72 and GB200 NVL72 are sold by the rack on reserved terms: 72 GPUs and 36 Grace CPUs in 18 compute trays, joined by 9 NVLink switch trays, in a liquid-cooled rack.
NVIDIA GB300 NVL72: 18 compute trays (2 Grace CPUs and 4 GPUs each), 9 NVLink switch trays (2 NVLink switch chips each), 2 management switches and 8 power shelves, with liquid-cooling manifolds and a bus bar at the rear. One compute tray and one switch tray are drawn pulled out.
KeyDWG 02
- 1. Power shelf · 8 per rack, six 5.5 kW supplies each
- 2. Management switch · 2 per rack
- 3. Compute tray · 18 × 1RU, 2 Grace CPUs and 4 GPUs each
- 4. NVLink switch tray · 9 × 1RU
- 5. NVLink switch chip · 2 per switch tray
- 6. Grace CPU · 2 per compute tray, 36 per rack
- 7. Blackwell Ultra GPU · 4 per compute tray, 72 per rack
- 8. ConnectX-8 SuperNIC, 800 Gb/s · 4 per compute tray
- 9. Liquid-cooling manifold · rear, supply and return
- 10. Bus bar · rear, fed by the power shelves
Spot: follow the market, or set your max.
Every GPU type has a spot price that moves with supply and demand. Follow it with no ceiling, or set the most you will pay per GPU-hour. You pay the spot price, not your max. If the price rises above your max, those nodes stop.
Drag the max price line, click the plot, or use the arrow keys.
- Max price
- $2.40 /GPU·hr
- Ran
- 51 h 15 m of 72 h
- Stopped
- 4×
- Resumed
- 4×
Blocked at start
Illustrative series, not Fantasti market data. A step line shows a spot price for one GPU type over 72 hours in 15-minute steps, between $1.91 and $2.72 per GPU-hour. With the max price at $2.40, the nodes run 51 h 15 m of 72 hours, stop 4 times when the price rises above the max, and resume 4 times when it falls back. You pay the spot price for each interval, not the max.
Computers your agents create and throw away.
Sandboxes are isolated CPU and GPU environments with a time-to-live and a network policy. An agent creates them through the API or the MCP server, runs untrusted code, and each one ends on schedule and stops billing.
fx.sandboxes.create(n=24, ttl="15m", network="closed") - Time-to-liveset per sandbox; it ends on schedule even if its agent is gone
- Network policyopen, restricted to an allowlist, or closed
- Meteringby the second, and it stops when the sandbox does
- Running
- 9 / 24
- Ended
- 15
- Billed
- 12,779 sandbox·s
- Running
- Ended
- TTL at 15:00
- Closed
- Restricted while running
One call creates 24 sandboxes with a 15-minute TTL. At 12 minutes, 9 are running and 15 have exited. Five more exit before 15 minutes, when the TTL stops the last 4. From 15 minutes on, nothing runs and nothing is billed. 4 sandboxes were switched from closed to restricted networking while running. Billed in total: 14,025 sandbox-seconds.
Built to be run by agents. The owner sets the rules.
Placement, pricing, capacity and support on Fantasti are built as desks of AI agents. Each agent has a name, a title, a lead it reports to and a list of decisions that go to senior review. The desks open with the first cohort. Until then the owner reads every request and answers it.
-
Early access
Tamsin
Support deskAI agent
- Works on
- How-to and status answers, and your ticket's progress
- Senior review
- Refunds and credits above $500
- Tier
- Front line
- Reports to
- Zainab · Head of Support · AI agent
-
Early access
Jonas
Technical supportAI agent
- Works on
- Diagnosis from your logs, metrics and placement records
- Senior review
- Anything on shared infrastructure
- Tier
- Senior
- Reports to
- Zainab · Head of Support · AI agent
-
Early access
Kenji
Solutions architectAI agent
- Works on
- Cluster designs, capacity plans and cost models
- Senior review
- Every design sent as a commitment
- Tier
- Orchestrator tier
- Reports to
- Idris · Head of Sales · AI agent
-
Private Preview
Leila
Capacity brokerAI agent
- Works on
- Placement and timing under your max price
- Senior review
- Capacity requests, changes to the Broker's strategy
- Tier
- Senior
- Reports to
- Maren · Orchestrator · AI agent
-
Early access
Hiro
Quota analystAI agent
- Works on
- Your account quota as you grow
- Senior review
- Increases above the Direct plan limits, any refusal
- Tier
- Senior
- Reports to
- Leila · Capacity broker · AI agent
-
Early access
Rafael
Risk and fraud officerAI agent
- Works on
- One rule set for sign-ups, payments and launches
- Senior review
- Suspension, closure, every appeal
- Tier
- Senior
- Reports to
- Maren · Orchestrator · AI agent
Every agent identifies itself as an AI in its first message and in its signature. All 10 teams and their roster
Front-line and technical agents decide everything inside their written playbook and limits. A high-stakes decision goes to senior review, by the orchestrator tier: refunds above $500, account suspension, capacity requests, security disclosures, playbook exceptions. The owner sets the rules and the hard limits, and performs the few acts only a legal person can perform. Every decision is logged with its inputs and its reasoning. Connections: The owner to Agents: Sets the rules; Desk to Senior review: High stakes; Agents to Decision record: Writes.
The owner sets the rules and the hard limits, andperforms the few acts only a legal person canperform.
- Money out
- Rule changes
- Stop
- Record
Front-line and technical agents
Everything inside their writtenplaybook and limits.
10 teamsBy the orchestrator tier
- Refunds above $500
- Account suspension
- Capacity requests
- Security disclosures
- Playbook exceptions
role · playbook · tier · inputsrule · reasoning · outcome · reviewEvery decision is logged with its inputs and its reasoning.
The owner sets the rules and the hardlimits, and performs the few acts only alegal person can perform.
- Money out
- Rule changes
- Stop
- Record
Front-line and technical agents
Everything inside their written playbookand limits.
10 teamsBy the orchestrator tier
- Refunds above $500
- Account suspension
- Capacity requests
- Security disclosures
- Playbook exceptions
role · playbook · tier · inputsrule · reasoning · outcome · reviewEvery decision is logged with its inputsand its reasoning.
- A high-stakes decision
- Sets the rules
Quota · Pay-as-you-go Early access
Your quota grows with your account.
Your quota grows automatically, on your payment history and your commitment. Up to the limits of the Direct plan there is no sales step. The Quota analyst raises your quota in steps and writes a decision record.
Hiro Quota analystAI agent
- At review
- quota 16peak 13headroom 3
- Now
- quota 24peak 18headroom 6
Decisiondec_4k9t…
Hiro Quota analystAI agent
raised · 16 → 24desk · inside the limits
Illustrative. Demand for H200 GPUs on the account acme rises in steps toward its quota of 16. A dashed review line sits under the quota. When demand reaches the review line, at a peak of 13, the quota is raised one step to 24 and a decision record is written. Demand passes the old quota later and never reaches a ceiling.
An agent proposes a change to a playbook, a price or a term. Senior review evaluates it before it takes effect.
In writingThe machine runs in the bay. What you sign is on paper.
One statement for all of it.
On-demand, spot and reserved usage from every product lands on one statement, split by cost center. Metered by the second, billed hourly.
Statement · Illustrative
| Cost center | Product | Listing | Pool | Metered GPU·s | Billed GPU·hr | Rate | Amount |
|---|---|---|---|---|---|---|---|
| cc-4102research | GPU Clusters, on-demand | H100 SXM5 | us-east | 21,024,000 | 5,840.00 | $5.40 | $31,536.00 |
| cc-4102research | GPU Instances, on-demand | H200 SXM | us-east | 1,843,200 | 512.00 | $5.94 | $3,041.28 |
| Subtotal cc-4102 | $34,577.28 | ||||||
| cc-2207evals | Batch Inference, spot | H100 SXM5 | us-central | 1,476,000 | 410.00 | $2.11 | $863.56 |
| cc-2207evals | Sandboxes | L40S | us-east | 301,968 | 83.88 | $1.86 | $156.02 |
| Subtotal cc-2207 | $1,019.58 | ||||||
| cc-3310serving | Serverless, on-demand | RTX PRO 6000 | us-central | 2,628,000 | 730.00 | $2.16 | $1,576.80 |
| Subtotal cc-3310 | $1,576.80 | ||||||
| Total | $37,173.66 | ||||||
| Cost center | Line | Billed GPU·hr | Rate | Amount |
|---|---|---|---|---|
| cc-4102 | GPU Clusters, reserved · B200 · rsv_7d2k | 11,680.00 | Order form | Per order form |
| account | Support plan · Direct | Monthly | Flat | $1,000.00 |
Support with a price list and a clock.
Every plan is designed so that an AI agent answers at any hour. Each plan states what it costs, its first response target and when an escalation gets its senior review.
Basic is included with every account. Paid plans are flat monthly fees on your Fantasti statement, with no percentage of spend.
| Plan | Price per month | First response AI agent, any hour | Senior review A second AI agent | Channels |
|---|---|---|---|---|
| BasicEveryone | Included | 15 min | Billing disputes and security reports only | Console and email |
| DeveloperIndividual developers | $29 | 10 min | Within 2 business days | Console and email |
| TeamStartups running production | $100 | 5 min | Within 1 business day | Console and email |
| DirectTeams with reserved or multi-node capacity | $1,000 | 5 min | Same business day | Console, email and a shared Slack channel |
| EnterpriseLarge organisations | Custom | Custom, in your order form | Custom, in your order form | Console, email and private Slack Connect |
First response is the time from opening a case to an AI agent's first reply on it. It is a target for response, not for resolution. Senior review is the time from an escalation to a reply from a second, more senior AI agent that has read the case. It is counted in business days.
Built in the open.
Every release is logged by product with its stage. Join a preview to use a feature early, vote on what comes next, and compare notes in the members' forum.
- 01 Changelog What changed, by product and stage. 2026-10-10 First cohort: requests open to companies Private Preview 2026-10-10 CPU instance and storage prices published Private Preview 2026-10-10 Billing terms: prepaid credit, pay-as-you-go and Enterprise Private Preview
- 02 Previews Use features before release.
- 03 Planned Feedback Vote. We decide, and we say why.
- 04 Planned Community The members' forum. Console account required.
- 05 Research GPU delivery over the network. Research in progress, no performance claims.
- Sheet
- 02 / 02
- Title
- In writing
- Reviewed
- 2026-10-10
Tell us what you need to run.
Companies can request a place in the first cohort, and no account is open yet. Choose a GPU, a count and a start date. We reply to a reviewed request, usually within one business day.
RequestFantasti Cloud Platform
Private PreviewReceived
request r_
Request received.
We reply to a reviewed request, usually within one business day.
You are on the waitlist.
Requests from personal email addresses join the waitlist for the next phase. We will write when a place opens.
Dev preview: no endpoint is configured, so nothing left this browser. Production keeps the form and says that nothing was sent.
- Sheet
- 01 / 01
- Prices reviewed
- 2026-10-10
Or read first
- GPU catalog Specifications, statuses and preview list prices.
- Pricing On-demand, spot and reserved, on one statement.
- Full request form Clusters, racks, regions and reserved terms.
Fantasti Cloud PlatformPrivate Preview
How an account is opened
-
You send a request through the form, from your company email address, and say what you want to run.
-
The owner reads every request for the first cohort. A request from a company domain is reviewed for fit. We reply to a reviewed request, usually within one business day.
-
While places remain, an approved request is invited by a sign-up link sent to that address. The sign-up link is tied to the email address of the request, is valid for 72 hours and works once.
- Valid for 72 hours
- Works once
- Tied to the email address of the request
-
Personal and free email addresses join the waitlist in this phase. When the cohort is full, new company requests join the waitlist too.