RunIntegrations Early access
Bring the scheduler you already use.
Every Fantasti cluster is yours, with its own Kubernetes API. Run Ray, SkyPilot or NVIDIA OSMO on it the way their docs describe, track runs in managed MLflow, and add applications from a catalog.
infra: k8s/fantasti-acme export KUBECONFIG="$HOME/.kube/fantasti-acme.yaml"
kubectl get nodes -L nvidia.com/gpu.product,nvidia.com/gpu.count Output · illustrative
NAME STATUS ROLES GPU.PRODUCT GPU.COUNT
gpu-node-01 Ready <none> NVIDIA-H200 8
gpu-node-02 Ready <none> NVIDIA-H200 8 §01 · On this page Kubernetes
helm repo add kuberay https://ray-project.github.io/kuberay-helm/
helm install kuberay-operator kuberay/kuberay-operator
# ray-values.yaml: a GPU worker group, min 0
helm install raycluster kuberay/ray-cluster -f ray-values.yaml §02 · On this page Ray
sky check kubernetes # finds your Fantasti context
sky gpus list --infra k8s/fantasti-acme # GPUs in your cluster
sky jobs launch -n sft-01 sft.yaml # recovers after node loss §03 · On this page SkyPilot
osmo pool list
osmo workflow submit train.yaml --pool fantasti-acme §04 · On this page NVIDIA OSMO
# MLFLOW_TRACKING_URI and credentials are already set
import mlflow
mlflow.set_experiment("sft-01")
with mlflow.start_run():
mlflow.log_params({"lr": 2e-5, "epochs": 3})
mlflow.log_metric("eval_loss", eval_loss)
mlflow.log_artifact("checkpoints/last/config.json") §05 · On this page MLflow
app = fx.apps.install("jupyterlab", cluster="acme-train", gpu="L40S")
print(app.url) §06 · On this page Applications
- Control plane
- your own, one per cluster
- GPU nodes
- dedicated · NVIDIA GPU labels
- Fabric
- one InfiniBand fabric per cluster
- Capacity
- on-demand · spot pools with max_price
Your cluster, your API server.
Each account gets an isolated cluster with its own control plane and dedicated GPU nodes. You receive a kubeconfig. Helm, operators and GitOps tools work unchanged, and kubectl get pods -A shows your workloads and nothing else.
- GPU labelsnodes carry the standard NVIDIA GPU product labels
- One fabricnodes in a cluster share one InfiniBand fabric
- Spot poolsadd spot node pools with a max price
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
cluster = fx.clusters.create(
name="sft-01",
gpu="H200:8", # 8-GPU nodes
nodes=2, # on one InfiniBand fabric
capacity=fantasti.Spot(max_price="follow"),
interfaces=["kubernetes", "skypilot"],
)
cluster.wait("ready")
cluster.kubeconfig.save("~/.kube/fantasti-acme.yaml") # Console: Clusters → acme → Connect → download kubeconfig
export KUBECONFIG="$HOME/.kube/fantasti-acme.yaml"
# your workloads and your cluster's system pods, nothing else
kubectl get pods -A
kubectl get nodes -L nvidia.com/gpu.product,nvidia.com/gpu.count Output · illustrative
NAME STATUS ROLES GPU.PRODUCT GPU.COUNT
gpu-node-01 Ready <none> NVIDIA-H200 8
gpu-node-02 Ready <none> NVIDIA-H200 8 {
"mcpServers": {
"fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
}
} Clusters tool group · planned
clusters_list
clusters_get
clusters_create asks for confirmation
clusters_delete asks for confirmation curl -X POST https://api.fantasti.ai/v1/clusters \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "acme-train",
"gpu": "H200:8",
"nodes": 2,
"capacity": { "mode": "spot", "max_price": "follow" },
"interfaces": ["kubernetes", "skypilot"]
}' Response · illustrative
{ "id": "cl_8f2k…", "status": "placing", "fabric": "single" } Planned Cluster tools for MCP clients such as coding agents, through the MCP server.
Ray on your GPUs.
Run Ray clusters with KubeRay on your Fantasti cluster, with GPU worker groups that scale from zero. Workspaces use the same Ray clusters, so code moves from notebook to job without changes.
# Plain Ray, nothing Fantasti-specific
import ray
ray.init() # connects to the Ray cluster on your Fantasti cluster
ds = ray.data.read_parquet("s3://acme-data/prompts/")
ds = ds.map_batches( # GPU workers scale up from 0
Embedder, concurrency=8, num_gpus=1, batch_size=256,
)
ds.write_parquet("s3://acme-data/embeddings/") helm repo add kuberay https://ray-project.github.io/kuberay-helm/
helm install kuberay-operator kuberay/kuberay-operator
# ray-values.yaml: a GPU worker group, min 0
helm install raycluster kuberay/ray-cluster -f ray-values.yaml Plain Ray code calls ray.init() and maps Embedder over a dataset with 8 workers of 1 GPU each. It connects to the Ray head that KubeRay runs on your Fantasti cluster. The GPU worker group starts at zero workers, scales to 8 while the job runs and returns to zero. Connections: Your code to Ray head; Ray head to GPU worker group.
ray.init() # plain Rayds.map_batches(Embedder,concurrency=8, num_gpus=1)- kind
- RayCluster
- runs on
- your cluster
- workers
- min 0 · max 8
- 1 GPU each
ray.init() # plain Rayds.map_batches(Embedder,concurrency=8, num_gpus=1)- kind
- RayCluster
- runs on
- your cluster
- workers
- min 0 · max 8
- 1 GPU each
ray.init() # plain Rayds.map_batches(Embedder,concurrency=8, num_gpus=1)- kind
- RayCluster
- runs on
- your cluster
- workers
- min 0 · max 8
- 1 GPU each
- Scales up
- Your job
1 Kubernetes API per cluster, and it is yours.
5 tools reach one Kubernetes API, which is the control plane of your own cluster: Kubernetes (early access), Ray (early access), SkyPilot (early access), NVIDIA OSMO (planned), Applications (planned). Your kubeconfig points at it, and kubectl get pods -A lists your workloads and your cluster's system pods, nothing else. The API schedules onto dedicated GPU nodes that share one InfiniBand fabric, one of them from a spot pool with a max price. Every node reports runs to managed MLflow. Connections: Kubernetes to Kubernetes API: kubeconfig; Ray to Kubernetes API; SkyPilot to Kubernetes API: kube context; NVIDIA OSMO to Kubernetes API; Applications to Kubernetes API; Kubernetes API to gpu-node-01; Kubernetes API to gpu-node-02; Kubernetes API to gpu-node-03: spot pool, max price; gpu-node-01 to gpu-node-02; gpu-node-02 to gpu-node-03; One InfiniBand fabric to Managed MLflow.
Early access
$ kubectl get pods -A- Helm
- operators
Early access
ray.init()- KubeRay
- from zero
Early access
infra: k8s/fantasti-acme- sky launch
- sky serve
Planned
$ osmo workflow submit- OSMO pools
Planned
fx.apps.install(…)- JupyterLab
- vLLM
- control plane
- yours, one per cluster
- kubeconfig
- fantasti-acme.yaml
- context
- fantasti-acme
- GPU nodes
- dedicated
$ kubectl get pods -ANAMESPACE NAME STATUSdefault kuberay-operator Runningdefault raycluster-head Runningkube-system coredns Running- API server
- 1
- GPU nodes
- 3
- GPUs
- 24
- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- tracking
- experiments, runs
- registry
- model versions
MLFLOW_TRACKING_URI- Workspaces
- Jobs
- Batch runs
Set on every node.
Early access
kubectlEarly access
ray.init()Early access
sky launchPlanned
osmoPlanned
fx.apps- control plane
- yours, one per cluster
- kubeconfig
- fantasti-acme.yaml
- context
- fantasti-acme
- GPU nodes
- dedicated
$ kubectl get pods -ANAMESPACE NAME STATUSdefault kuberay-operator Runningdefault raycluster-head Runningkube-system coredns Running- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- tracking
- experiments, runs
- registry
- model versions
MLFLOW_TRACKING_URI- Workspaces
- Jobs
- Batch runs
Set on every node.
- control plane
- yours, one per cluster
- kubeconfig
- fantasti-acme.yaml
- context
- fantasti-acme
- GPU nodes
- dedicated
$ kubectl get pods -ANAMESPACE NAME STATUSdefault kuberay-operator Runningdefault raycluster-head Runningkube-system coredns Running- 8 GPUs
- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- 8 GPUs
- gpu.product
- NVIDIA-H200
- gpu.count
- 8
- tracking
- experiments, runs
- registry
- model versions
MLFLOW_TRACKING_URI
- Your kubeconfig
- Reports runs
- Planned route
- Planned
Bring your SkyPilot YAML.
Fantasti issues a Kubernetes context for your cluster. Point SkyPilot at it and sky launch, sky jobs launch and sky serve work as they do anywhere else, with failover across the GPU types you list.
infra: k8s/fantasti-acme SkyPilot’s own “workspaces” are team permissions. Fantasti Workspaces are GPU development environments; they are different things.
# sft.yaml: "fantasti-acme" is the kube context for your isolated cluster
name: sft-01
num_nodes: 2
resources:
infra: k8s/fantasti-acme
accelerators: {H200:8, B200:8} # either works; SkyPilot takes what is free
job_recovery:
strategy: FAILOVER
max_restarts_on_errors: 2
workdir: .
setup: pip install -r requirements.txt
run: |
MASTER=$(echo "$SKYPILOT_NODE_IPS" | head -n1)
torchrun --nnodes=$SKYPILOT_NUM_NODES --node_rank=$SKYPILOT_NODE_RANK \
--nproc_per_node=$SKYPILOT_NUM_GPUS_PER_NODE \
--master_addr=$MASTER --master_port=29500 train.py --resume sky check kubernetes # finds your Fantasti context
sky gpus list --infra k8s/fantasti-acme # GPUs in your cluster
sky jobs launch -n sft-01 sft.yaml # recovers after node loss NVIDIA OSMO for physical-AI workflows. Planned
Use OSMO on your Fantasti cluster to run simulation, synthetic-data and training workflows as one pipeline. Fantasti installs it and exposes your GPUs as OSMO pools.
# train.yaml: an OSMO workflow on your Fantasti cluster's GPU pool
workflow:
name: policy-train
resources:
gpu_task:
cpu: 16
gpu: 1
memory: 64Gi
storage: 100Gi
tasks:
- name: train
image: nvcr.io/nvidia/pytorch:24.01-py3
command: ["bash", "-c", "python train.py --epochs 3"]
resource: gpu_task osmo pool list
osmo workflow submit train.yaml --pool fantasti-acme Three workflows, simulation, synthetic data and training, run in order as one pipeline. NVIDIA OSMO, installed by Fantasti on your cluster, submits their tasks to one pool named fantasti-acme, which is the GPU nodes of your cluster. The task train of the workflow policy-train holds 1 of the 8 GPUs of the first node. Connections: Simulation to Synthetic data; Synthetic data to Training; Simulation to NVIDIA OSMO; Synthetic data to NVIDIA OSMO; Training to NVIDIA OSMO; NVIDIA OSMO to gpu-node-01: submits tasks.
fantasti-acmeYour GPU nodesInstalled by Fantasti on your cluster.
- workflow
- policy-train
- pool
- fantasti-acme
- task train
- free
fantasti-acmeInstalled by Fantasti on your cluster.
- workflow
- policy-train
- pool
- fantasti-acme
- task train
- free
- One pipeline
- Planned
Managed MLflow.
Managed MLflow tracking and model registry, wired into Workspaces, Flows, Batch Inference, Serverless jobs and Clusters.
A workload on any of your nodes, in Workspaces, jobs or batch runs, already has MLFLOW_TRACKING_URI and credentials set. It logs the parameter lr and the metric eval_loss to the experiment sft-01, which holds 3 runs. Run 03 is registered as version 3 of the model sft-01. Connections: Any workload to sft-01: logs; sft-01 to sft-01: registers.
- Workspaces
- Jobs
- Batch runs
- MLFLOW_TRACKING_URI
- set
- credentials
- set
mlflow.log_params({"lr": 2e-5})mlflow.log_metric("eval_loss", …)run lr eval_loss03 2e-5 0.41202 5e-5 0.44701 1e-4 0.503- artifacts
- checkpoints/last
v3 run 03v2 run 02v1 run 01
- Workspaces
- Jobs
- Batch runs
- MLFLOW_TRACKING_URI
- set
- credentials
- set
mlflow.log_params({"lr": 2e-5})mlflow.log_metric("eval_loss", …)run lr eval_loss03 2e-5 0.41202 5e-5 0.44701 1e-4 0.503- artifacts
- checkpoints/last
v3 run 03v2 run 02v1 run 01
- Workspaces
- Jobs
- Batch runs
- MLFLOW_TRACKING_URI
- set
- credentials
- set
mlflow.log_params({"lr": 2e-5})mlflow.log_metric("eval_loss", …)run lr eval_loss03 2e-5 0.41202 5e-5 0.44701 1e-4 0.503- artifacts
- checkpoints/last
v3 run 03v2 run 02v1 run 01
- Registered
- Logged by your code
# MLFLOW_TRACKING_URI and credentials are already set
import mlflow
mlflow.set_experiment("sft-01")
with mlflow.start_run():
mlflow.log_params({"lr": 2e-5, "epochs": 3})
mlflow.log_metric("eval_loss", eval_loss)
mlflow.log_artifact("checkpoints/last/config.json") -
01Tracking
Experiments and runs
Log parameters, metrics and artifacts from any workload.
-
02Registry
Model versions
Register, version and promote models.
-
03Wired in
Set on every node
Workspaces, jobs and batch runs get
MLFLOW_TRACKING_URIand credentials.
Applications, one click from your GPUs. Planned
| Category | Applications | Managed app | On your cluster | Runs as |
|---|---|---|---|---|
| Notebooks and IDEs |
| Yes | Yes | Managed app or on your cluster |
| Inference servers |
| No | Yes | On your cluster |
| Vector databases |
| Yes | Yes | Managed app or on your cluster |
| Interfaces |
| Yes | No | Managed app |
| Schedulers |
| No | Yes | On your cluster |
app = fx.apps.install("jupyterlab", cluster="acme-train", gpu="L40S")
print(app.url) Questions and answers.
Do I need to change my Ray or SkyPilot code?
Can my scheduler run on spot capacity?
Which integrations are in early access?
What about Metaflow?
- Sheet
- 01
- Title
- Integrations · applications and answers
- Reviewed
- 2026-10-09
Keep your scheduler. Change where it runs.
Integrations is in early access. Request access and tell us which scheduler you run.