Fantasti Cloud Platform Private Preview
Every layer of a GPU cloud, on one account.
The Fantasti Cloud Platform is compute, development environments, ways to run work, storage and networking, and the tools to operate them. Every product shares one API, one set of policies, one quota system and one statement.
fx = fantasti.Client() - account
- acme
- products
- 20 · one client Products: §01 Compute: 5 (4 Private Preview, 1 By request). Develop: 3 (2 Early access, 1 Planned). Data and network: 2 (2 Private Preview). Run: 3 (1 Private Preview, 2 Early access). Operate: 4 (1 Private Preview, 2 Early access, 1 Planned). Access: 3 (2 Private Preview, 1 Early access).
- policy
- research.yaml · applies to every request Shared controls: §06
- capacity
- on-demand · spot · reserved Capacity: §03
- quota
- reviewed by the Orchestrator Quota: Orchestrator
- statement
- one · split by cost center One bill: Pricing
- Interfaces
- Python · REST · CLI · MCP
- NVIDIA GPUs
- 6 models · 2 rack-scale systems
- Metering
- per second · one statement
- Stage
- Private Preview · requests open
What is on the platform.
-
015 products
Compute
GPUs and CPUs as instances, clusters and racks.
- GPU Instances Virtual machines with one or eight GPUs.
fx.instances.create(gpu="H100:1") - CPU Instances General-purpose machines beside your GPUs.
fx.instances.create(cpu=32) - GPU Clusters 8-GPU nodes on one InfiniBand fabric.
fx.clusters.create(gpu="B200:8", nodes=16) - Rack-scale By request GB300 NVL72 and GB200 NVL72 racks, reserved.
capacity=Reserved("rsv_7d2k") - Spot Market-priced GPUs with a max price you set.
capacity=Spot(max_price=3.10)
- GPU Instances Virtual machines with one or eight GPUs.
-
023 products
Develop
Environments people and agents work in.
- Workspaces Early access GPU dev environments that scale to Ray clusters.
fx.workspaces.create(gpu="H100:1") - Sandboxes Early access Short-lived isolated environments for AI agents.
fx.sandboxes.create(ttl="15m") - Flows Planned Python workflows from experiment to production, on GPUs.
@fantasti(gpu="H200:8", capacity="spot")
- Workspaces Early access GPU dev environments that scale to Ray clusters.
-
032 products
Data and network
Storage and networking around your compute.
-
043 products
Run
Ways to run work without managing machines.
- Batch Inference Offline inference at market prices, with checkpoints.
fx.batch.create(input="s3://…") - Serverless Early access Run a container as a job or endpoint.
fx.jobs.create(image=…, gpu="H200:8") - Integrations Early access Kubernetes, Ray, SkyPilot and NVIDIA OSMO on your cluster.
infra: k8s/fantasti-acme
- Batch Inference Offline inference at market prices, with checkpoints.
-
054 products
Operate
The control plane and the shared tools.
- Orchestrator Policy, placement, pricing and recovery for every workload.
placement: { fabric: infiniband } - Observability Early access GPU metrics, logs and traces, with dashboards.
fx.metrics.query("…") - MLflow Early access Managed experiment tracking and model registry.
MLFLOW_TRACKING_URI - Applications Planned Notebooks, inference servers and databases, one click.
fx.apps.install("jupyterlab")
- Orchestrator Policy, placement, pricing and recovery for every workload.
-
3 products
Access
One console, one API, one bill.
Capacity underneath.
A cluster that is yours.
A capacity layer of NVIDIA GPU nodes underneath. On it, an isolated cluster for each account, running the tools you already use. You work through one API, one console and one bill.
Illustrative. Six products reach the platform through one API: GPU Clusters (Private Preview), Workspaces (Early access), Sandboxes (Early access), Flows (Planned), Batch Inference (Private Preview), Serverless (Early access). The figure follows one call, fx.clusters.create(gpu="H200:8", nodes=2). It arrives as the request POST /v1/clusters. The Fantasti Orchestrator checks it at five gates, each with the value the request carries: GPU H200:8 × 2, fabric infiniband, region us-*, capacity spot, max price ≤ 3.10. It is placed on the capacity layer, on us-east, fabric a, as cluster cl_8f2k: 2 nodes, 16 dedicated GPUs, one InfiniBand fabric and a control plane of your own. us-east, fabric b is passed over: 8 free < 16. The same layer holds rack-scale systems (72 GPUs in one NVLink domain, by request), and CPU Instances, Storage, Networking, Managed services. Usage is metered per second onto one statement for account acme, on the cost center research-llm. Connections: Workspaces to POST /v1/clusters; Sandboxes to POST /v1/clusters; Flows to POST /v1/clusters: One API; Batch Inference to POST /v1/clusters; Serverless to POST /v1/clusters; GPU Clusters to POST /v1/clusters; POST /v1/clusters to H200:8 × 2; H200:8 × 2 to infiniband; infiniband to us-*; us-* to spot; spot to ≤ 3.10; ≤ 3.10 to cl_8f2k: Placed; cl_8f2k to acme: Metered.
fx.clusters.create(gpu="H200:8",nodes=2)Private Previewfx.workspaces.create(gpu="H100:1")Early accessfx.sandboxes.create(ttl="15m")Early access@fantasti(gpu="H200:8",capacity="spot")Plannedfx.batch.create(input="s3://…")Private Previewfx.jobs.create(image=…,gpu="H200:8")Early access{ "name": "sft-01","gpu": "H200:8", "nodes": 2,"capacity": { "mode": "spot","max_price": "3.10" } }8 GPUs per node
- nodes
- 2 × H200:8
- fabric
- one · InfiniBand
- control
- your own plane
- state
- running
16 GPUs, dedicated- CPU Instances
- Storage
- Networking
- Managed services
Same API, console and bill.
- GB300 NVL72
- GB200 NVL72
72 GPUs in one NVLink domain.On reserved terms.
- Free 8Needs 16
8 free < 16One fabric per job.
- line
- cl_8f2k · GPU Clusters
- capacity
- spot ≤ 3.10
- cost center
- research-llm
- statement
- one · split by cost center
- metered
- per second
Every product on one statement, split by cost center.
fx.clusters.create(gpu="H200:8",nodes=2)Private Previewfx.workspaces.create(gpu="H100:1")Early accessfx.sandboxes.create(ttl="15m")Early access@fantasti(gpu="H200:8",capacity="spot")Plannedfx.batch.create(input="s3://…")Private Previewfx.jobs.create(image=…,gpu="H200:8")Early access{ "name": "sft-01","gpu": "H200:8", "nodes": 2,"capacity": { "mode": "spot","max_price": "3.10" } }# the API answers{ "id": "cl_8f2k…","status": "placing","fabric": "single" }8 GPUs per node
- nodes
- 2 × H200:8
- fabric
- one · InfiniBand
- control
- your own plane
- state
- running
Your own control plane anddedicated GPU nodes. No otheraccount runs on them.
16 GPUs, dedicated- GB300 NVL72
- GB200 NVL72
72 GPUs in one NVLink domain.On reserved terms.
- Free 8Needs 16
8 free < 16One fabric per job.
- CPU Instances
- Storage
- Networking
- Managed services
Same API, console and bill.
- line
- cl_8f2k · GPU Clusters
- capacity
- spot ≤ 3.10
- cost center
- research-llm
- statement
- one · split by cost center
- metered
- per second
- GPU Clusters
- Private Preview
- Workspaces
- Early access
- Sandboxes
- Early access
- Flows
- Planned
- Batch Inference
- Private Preview
- Serverless
- Early access
fx.clusters.create(gpu="H200:8",nodes=2){ "name": "sft-01","gpu": "H200:8", "nodes": 2,"capacity": { "mode": "spot","max_price": "3.10" } }- GPU
- H200:8 × 2
- fabric
- infiniband
- region
- us-*
- capacity
- spot
- max price
- ≤ 3.10
8 GPUs per node
- nodes
- 2 × H200:8
- fabric
- one · InfiniBand
- control
- your own plane
- state
- running
16 GPUs- Free 8Needs 16
8 free < 16One fabric per job.
- GB300 NVL72
- GB200 NVL72
72 GPUs in one NVLink domain.On reserved terms.
- CPU Instances
- Storage
- Networking
- Managed services
Same API, console and bill.
- line
- cl_8f2k · GPU Clusters
- capacity
- spot ≤ 3.10
- cost center
- research-llm
- statement
- one · split by cost center
- metered
- per second
- This request
- Every other product
- Planned
- Passed over
The capacity layer comes from infrastructure providers. Fantasti builds the layer you use. You sign with Fantasti, call Fantasti and pay Fantasti.
-
You One console, one API, one bill.
-
Fantasti Builds the layer you use.
-
The capacity layer NVIDIA GPU capacity, storage and networks.
Pick the right product.
The Orchestrator is not another runtime. It places, limits and bills all of these.
Spot is a capacity mode, not a runtime. It applies to every column that says so.
| Question | GPU Instances | GPU Clusters | Workspaces Early access | Sandboxes Early access | Flows Planned | Batch Inference | Serverless Early access |
|---|---|---|---|---|---|---|---|
| You bring | An image, or your own stack | A scheduler, or nothing but SSH | Your editor and code | An agent or test harness | A Python flow | A model and a dataset | A container |
| Fantasti runs | One machine with 1 or 8 GPUs | Many 8-GPU nodes on one fabric | A dev environment that scales to a cluster | Short-lived isolated environments | Each step, in order, on the GPU it names | A sharded inference run | One job or one endpoint |
| Ends | When you stop it or its term ends | When you stop it or its term ends | When you stop it | At its time-to-live | When the last step does; then on its next trigger | When every shard is done | When the job exits or you stop the endpoint |
| Keeps | Its volumes | Its volumes and shared filesystem | Your files | Nothing; copy results out | Every run, versioned | Outputs and checkpoints | Nothing on the container disk |
| Capacity | On-demand, spot, reserved | On-demand, spot workers, reserved | On-demand head; workers on-demand or spot | On-demand | Per step: on-demand or spot | Spot or on-demand | On-demand or spot |
| Use it when | You need a GPU and root | The job spans nodes | You are writing or debugging | Code is untrusted or disposable | The work repeats and must be reproducible | The job is one model over one dataset | The job is one container |
- Sheet
- 01 / 02
- Title
- Capacity and product comparison
- Reviewed
- 2026-10-10
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
cluster = fx.clusters.create(
name="sft-01",
gpu="H200:8", # 8-GPU nodes
nodes=2, # placed on one InfiniBand fabric
capacity=fantasti.Spot(max_price=3.10), # your ceiling, USD per GPU-hour
interfaces=["kubernetes", "skypilot"],
)
cluster.wait("ready")
cluster.kubeconfig.save("~/.kube/fantasti-acme.yaml") $ fantasti cluster create sft-01 --gpu H200:8 --nodes 2 --spot --max-price 3.10 Output · illustrative
request H200:8 × 2 · one InfiniBand fabric · spot ≤ 3.10
pools 4 considered
placed cl_8f2k… · us-east · fabric a
state running · kubeconfig saved {
"mcpServers": {
"fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
}
} Tools · early access
catalog gpus_list
sandboxes sandboxes_create · sandboxes_exec · sandboxes_read_file · sandboxes_write_file · sandboxes_delete
workspaces workspaces_list · workspaces_start · workspaces_stop
jobs jobs_create · jobs_logs
usage usage_get curl -X POST https://api.fantasti.ai/v1/clusters \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "acme-train",
"gpu": "H200:8",
"nodes": 2,
"capacity": { "mode": "spot", "max_price": "3.10", "currency": "USD", "unit": "gpu_hour" },
"interfaces": ["kubernetes", "skypilot"]
}'
# { "id": "cl_8f2k…", "status": "placing", "fabric": "single" } # policy.yaml · applies to every request from team "research"
team: research
placement:
gpus: [H100, H200, B200]
fabric: infiniband # multi-node jobs stay on one fabric
regions: ["us-*"]
capacity:
allow: [on_demand, spot, reserved]
spot:
max_price: { H100: 2.40, H200: 3.10 } # USD per GPU-hour, your ceilings
limits:
gpus: { H200: 32 }
sandboxes: { concurrent: 200, max_ttl: 60m }
spend: { monthly_usd: 40000 }
cost_center: research-llm One package, one key.
Every product shares the client, the authentication and the error model. Every object reports where it was placed and what it costs.
- Python
- import fantasti
- CLI
- fantasti <noun> <verb>
- REST
- JSON over HTTPS
- MCP
- remote server
Set once, applied everywhere.
Policies written in the Orchestrator apply to every request, whichever product makes it and whoever makes it: a person, a pipeline or an agent.
A request reaches the Fantasti Orchestrator from a person, a pipeline or an agent, for any product: a person through CLI, here for GPU Clusters; a pipeline through Python, here for Batch Inference; an agent through MCP, here for Sandboxes. Every request passes the same five shared controls, in this order. 1, Identity: Single sign-on and role-based access. Set in console · access. 2, Placement policy: GPU, interconnect, region, capacity mode and max price. Set in policy.yaml · placement, capacity. 3, Quotas and limits: A pay-as-you-go quota that grows with you, and team limits your admins set. Set in policy.yaml · limits. 4, Cost centers: Every statement line is attributed. Set in policy.yaml · cost_center. 5, One statement: Every product, metered by the second. Set in statement · cost centers. Connections: A person to Fantasti Orchestrator; A pipeline to Fantasti Orchestrator: Every request; An agent to Fantasti Orchestrator; Identity to Placement policy; Placement policy to Quotas and limits; Quotas and limits to Cost centers; Cost centers to One statement.
$ fantasti cluster create sft-01--gpu H200:8 --nodes 2 --spot --max-price 3.10fx.batch.create(input="s3://…")# reads FANTASTI_API_KEYsandboxes_create# OAuth, or a scoped key- person
- SSO
- pipeline
- API key
- agent
- scoped key
Single sign-on androle-based access.
# policy.yamlplacement:gpus: [H100, H200, B200]fabric: infinibandregions: ["us-*"]capacity:allow: [on_demand, spot, reserved]spot:max_price: { H100: 2.40, H200: 3.10 }GPU, interconnect, region, capacity mode andmax price.
# policy.yamllimits:gpus: { H200: 32 }spend: { monthly_usd: 40000 }A pay-as-you-go quota that grows withyou, and team limits your admins set.
# policy.yamlcost_center: research-llm- cluster
- cl_8f2k
- billed to
- research-llm
Every statement line isattributed.
- metered
- per second
- split
- cost center
Every product, meteredby the second.
$ fantasti cluster create sft-01--gpu H200:8 --nodes 2 --spot--max-price 3.10fx.batch.create(input="s3://…")# reads FANTASTI_API_KEYsandboxes_create# OAuth, or a scoped key- person
- SSO
- pipeline
- API key
- agent
- scoped key
Single sign-on and role-basedaccess.
# policy.yamlplacement:gpus: [H100, H200, B200]fabric: infinibandregions: ["us-*"]capacity:allow: [on_demand, spot, reserved]spot:max_price: { H100: 2.40, H200: 3.10 }GPU, interconnect, region, capacity mode and max price.
# policy.yamllimits:gpus: { H200: 32 }spend: { monthly_usd: 40000 }A pay-as-you-go quota that grows withyou, and team limits your admins set.
# policy.yamlcost_center: research-llm- cluster
- cl_8f2k
- billed to
- research-llm
Every statement line isattributed.
- metered
- per second
- split
- cost center
Every product, metered bythe second.
- A person
- CLI · GPU Clusters
- A pipeline
- Python · Batch Inference
- An agent
- MCP · Sandboxes
- person
- SSO
- pipeline
- API key
- agent
- scoped key
Single sign-on and role-based access.
# policy.yamlplacement:gpus: [H100, H200, B200]fabric: infinibandregions: ["us-*"]capacity:allow: [on_demand, spot, reserved]spot:max_price: { H100: 2.40, H200: 3.10 }GPU, interconnect, region, capacity mode and maxprice.
# policy.yamllimits:gpus: { H200: 32 }spend: { monthly_usd: 40000 }A pay-as-you-go quota that grows with you, andteam limits your admins set.
# policy.yamlcost_center: research-llm- cluster
- cl_8f2k
- billed to
- research-llm
Every statement line is attributed.
- metered
- per second
- split
- cost center
Every product, metered by the second.
- The path of every request
- A request
A request reaches the Fantasti Orchestrator from a person, a pipeline or an agent, for any product: a person through CLI, here for GPU Clusters; a pipeline through Python, here for Batch Inference; an agent through MCP, here for Sandboxes. Every request passes the same five shared controls, in this order. 1, Identity: Single sign-on and role-based access. Set in console · access. 2, Placement policy: GPU, interconnect, region, capacity mode and max price. Set in policy.yaml · placement, capacity. 3, Quotas and limits: A pay-as-you-go quota that grows with you, and team limits your admins set. Set in policy.yaml · limits. 4, Cost centers: Every statement line is attributed. Set in policy.yaml · cost_center. 5, One statement: Every product, metered by the second. Set in statement · cost centers. Connections: Every request to Fantasti Orchestrator; Identity to Placement policy; Placement policy to Quotas and limits; Quotas and limits to Cost centers; Cost centers to One statement.
- A person
- CLI · GPU Clusters
- A pipeline
- Python · Batch Inference
- An agent
- MCP · Sandboxes
- person
- SSO
- pipeline
- API key
- agent
- scoped key
Single sign-on and role-basedaccess.
# policy.yamlplacement:gpus: [H100, H200, B200]fabric: infinibandregions: ["us-*"]capacity:allow: [on_demand, spot, reserved]spot:max_price: { H100: 2.40, H200: 3.10 }GPU, interconnect, region, capacity mode and maxprice.
# policy.yamllimits:gpus: { H200: 32 }spend: { monthly_usd: 40000 }A pay-as-you-go quota that grows with you,and team limits your admins set.
# policy.yamlcost_center: research-llm- cluster
- cl_8f2k
- billed to
- research-llm
Every statement line is attributed.
- metered
- per second
- split
- cost center
Every product, metered by the second.
- The path of every request
- A request
Built to be run by agents.
Under written rules.
The platform's volume work is built for AI agents: answering, diagnosing, placing and pricing. Each has one job and a limit on what it may decide. High-stakes decisions pass a second, senior review. The owner sets the rules.
- Agents
- 13 roles · 3 planned
- Decide
- act · ask · hand off
- Levels
- desk · senior review · the owner
Raise a quota: the Quota analyst reads your payment history and commitment, the risk rules check the increase, and above the limits of the Direct plan it goes to senior review.
Three levels. 01 Desk: front-line and technical agents in 10 teams (Sales, Support, Engineering, Security, Review, Legal, Monitoring, Risk, Capacity, Release) decide everything inside their written playbook and limits. 02 Senior review, by the orchestrator tier, takes the high-stakes decisions: Refunds and credits above $500; Suspending or terminating an account; Reservations and capacity requests to the capacity layer; Security disclosures; Exceptions to a playbook. The reviewer is a different agent from the one that proposed the decision, and it records its reasoning. Everything else, agents decide inside their playbooks and log. 03 The owner: The rules and limits themselves, and the few acts only a legal person can perform. Signs contracts and legal documents for Fantasti. Owns the bank and payment accounts and passes their identity checks. Answers legal process and regulators, and files taxes. Sets and changes the hard limits and the decision rights. The levels work under the written rules, set by the owner: Catalog and prices (what a customer pays), Your policy (quotas, caps, ceilings), Risk rule set (allow, verify, hold, block), Playbooks (versioned and reviewed). Every decision is logged in a decision record: role · playbook · tier · inputs · rule · reasoning · outcome · review. Connections: Desk to Senior review: Hand off; Senior review to The owner: Digest; The owner to Written rules: Sets; Written rules to Decision record: Writes.
Everything inside their written playbook and limits.
- Sales
- Support
- Engineering
- Security
- Review
- Legal
- Monitoring
- Risk
- Capacity
- Release
- Anouk
- Front desk · AI agent
- Tamsin
- Support desk · AI agent
- Kofi
- QA reviewer · AI agent
- Hiro
- Quota analyst · AI agent
- Rafael
- Risk and fraud officer · AI agent
- Rafael
- Risk and fraud officer · AI agent
- Kofi
- QA reviewer · AI agent
- Freya
- Release manager · AI agent
- Refunds above $500
- Account suspension
- Capacity requests
- Security disclosures
- Playbook exceptions
The reviewer is a different agent from the one that proposed the decision, and it recordsits reasoning.
- None
- decided at the desk
- Ingrid
- Head of Review · AI agent
- Ingrid
- Head of Review · AI agent
- Ingrid
- Head of Review · AI agent
The rules and limits themselves, and the few acts only a legal person can perform.
- Contracts and legal documents
- Bank and payment accounts
- Legal process and regulators
- Hard limits and decision rights
What a customer pays
Quotas, caps, ceilings
Allow, verify, hold, block
Versioned and reviewed
- role
- front-desk, support-desk, qa-reviewer
- role
- quota-analyst, risk-officer
- role
- risk-officer
- role
- qa-reviewer, release-manager
- rule
- catalog and prices, playbooks
- rule
- your policy, risk rule set
- rule
- risk rule set
- rule
- playbooks
- tier
- front-line, senior
- tier
- senior
- tier
- senior
- tier
- senior, orchestrator-tier
- review
- none
- review
- senior review
- review
- senior review
- review
- senior review
Every decision is logged: role · playbook · tier · inputs · rule · reasoning · outcome ·review.
Everything inside their written playbookand limits.
- Sales
- Support
- Engineering
- Security
- Review
- Legal
- Monitoring
- Risk
- Capacity
- Release
- Anouk
- AI agent
- Front desk
- Tamsin
- AI agent
- Support desk
- Kofi
- AI agent
- QA reviewer
- Hiro
- AI agent
- Quota analyst
- Rafael
- AI agent
- Risk and fraud officer
- Rafael
- AI agent
- Risk and fraud officer
- Kofi
- AI agent
- QA reviewer
- Freya
- AI agent
- Release manager
- Refunds above $500
- Account suspension
- Capacity requests
- Security disclosures
- Playbook exceptions
The reviewer is a different agent from theone that proposed the decision, and itrecords its reasoning.
- None
- decided at the desk
- Ingrid
- Head of Review · AI agent
- Ingrid
- Head of Review · AI agent
- Ingrid
- Head of Review · AI agent
The rules and limits themselves, and thefew acts only a legal person can perform.
- Contracts and legal documents
- Bank and payment accounts
- Legal process and regulators
- Hard limits and decision rights
- Catalog and prices
- Your policy
- Risk rule set
- Playbooks
- role
- front-desk, support-desk,
- qa-reviewer
- role
- quota-analyst, risk-officer
- role
- risk-officer
- role
- qa-reviewer, release-manager
- rule
- catalog and prices, playbooks
- rule
- your policy, risk rule set
- rule
- risk rule set
- rule
- playbooks
- tier
- front-line, senior
- tier
- senior
- tier
- senior
- tier
- senior, orchestrator-tier
- review
- none
- review
- senior review
- review
- senior review
- review
- senior review
Every decision is logged: role · playbook ·tier · inputs · rule · reasoning · outcome· review.
- The traced decision
- Works under the rules
- Digest to the owner
- Goes to senior review
Every product carries a stage.
-
Private Preview 10 products
Opens with the first cohort. Access by request.
-
Early access 7 products
Being built for the first cohort. Details may change.
- Sheet
- 02 / 02
- Title
- Release stages
- Reviewed
- 2026-10-10