RunIntegrations Early access

Bring the scheduler you already use.

Every Fantasti cluster is yours, with its own Kubernetes API. Run Ray, SkyPilot or NVIDIA OSMO on it the way their docs describe, track runs in managed MLflow, and add applications from a catalog.

infra: k8s/fantasti-acme
Being built for the first cohort. NVIDIA OSMO and Applications are planned.
TAB 01Interfaces to one cluster Source · Fantasti
Metaflow Runs as Fantasti Flows. Planned
CLI
Preview API · subject to change
export KUBECONFIG="$HOME/.kube/fantasti-acme.yaml"

kubectl get nodes -L nvidia.com/gpu.product,nvidia.com/gpu.count

Output · illustrative

NAME          STATUS  ROLES   GPU.PRODUCT  GPU.COUNT
gpu-node-01   Ready   <none>  NVIDIA-H200  8
gpu-node-02   Ready   <none>  NVIDIA-H200  8

§01 · On this page Kubernetes

CLI
Preview API · subject to change
helm repo add kuberay https://ray-project.github.io/kuberay-helm/
helm install kuberay-operator kuberay/kuberay-operator

# ray-values.yaml: a GPU worker group, min 0
helm install raycluster kuberay/ray-cluster -f ray-values.yaml

§02 · On this page Ray

CLI
Preview API · subject to change
sky check kubernetes                      # finds your Fantasti context
sky gpus list --infra k8s/fantasti-acme   # GPUs in your cluster
sky jobs launch -n sft-01 sft.yaml     # recovers after node loss

§03 · On this page SkyPilot

CLI
Preview API · subject to change
osmo pool list
osmo workflow submit train.yaml --pool fantasti-acme

§04 · On this page NVIDIA OSMO

Python
Preview API · subject to change
# MLFLOW_TRACKING_URI and credentials are already set
import mlflow

mlflow.set_experiment("sft-01")
with mlflow.start_run():
    mlflow.log_params({"lr": 2e-5, "epochs": 3})
    mlflow.log_metric("eval_loss", eval_loss)
    mlflow.log_artifact("checkpoints/last/config.json")

§05 · On this page MLflow

Python
Preview API · subject to change
app = fx.apps.install("jupyterlab", cluster="acme-train", gpu="L40S")
print(app.url)

§06 · On this page Applications

Control plane
your own, one per cluster
GPU nodes
dedicated · NVIDIA GPU labels
Fabric
one InfiniBand fabric per cluster
Capacity
on-demand · spot pools with max_price
§01 Kubernetes Early access

Your cluster, your API server.

Each account gets an isolated cluster with its own control plane and dedicated GPU nodes. You receive a kubeconfig. Helm, operators and GitOps tools work unchanged, and kubectl get pods -A shows your workloads and nothing else.

  • GPU labelsnodes carry the standard NVIDIA GPU product labels
  • One fabricnodes in a cluster share one InfiniBand fabric
  • Spot poolsadd spot node pools with a max price
Preview API · subject to change
import fantasti

fx = fantasti.Client()                # reads FANTASTI_API_KEY

cluster = fx.clusters.create(
    name="sft-01",
    gpu="H200:8",                     # 8-GPU nodes
    nodes=2,                          # on one InfiniBand fabric
    capacity=fantasti.Spot(max_price="follow"),
    interfaces=["kubernetes", "skypilot"],
)
cluster.wait("ready")
cluster.kubeconfig.save("~/.kube/fantasti-acme.yaml")
# Console: Clusters → acme → Connect → download kubeconfig
export KUBECONFIG="$HOME/.kube/fantasti-acme.yaml"

# your workloads and your cluster's system pods, nothing else
kubectl get pods -A

kubectl get nodes -L nvidia.com/gpu.product,nvidia.com/gpu.count

Output · illustrative

NAME          STATUS  ROLES   GPU.PRODUCT  GPU.COUNT
gpu-node-01   Ready   <none>  NVIDIA-H200  8
gpu-node-02   Ready   <none>  NVIDIA-H200  8
{
  "mcpServers": {
    "fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
  }
}

Clusters tool group · planned

clusters_list
clusters_get
clusters_create    asks for confirmation
clusters_delete    asks for confirmation
curl -X POST https://api.fantasti.ai/v1/clusters \
  -H "Authorization: Bearer $FANTASTI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "name": "acme-train",
        "gpu": "H200:8",
        "nodes": 2,
        "capacity": { "mode": "spot", "max_price": "follow" },
        "interfaces": ["kubernetes", "skypilot"]
      }'

Response · illustrative

{ "id": "cl_8f2k…", "status": "placing", "fabric": "single" }

Planned Cluster tools for MCP clients such as coding agents, through the MCP server.

§02 Ray Early access

Ray on your GPUs.

Run Ray clusters with KubeRay on your Fantasti cluster, with GPU worker groups that scale from zero. Workspaces use the same Ray clusters, so code moves from notebook to job without changes.

Preview API · subject to change
# Plain Ray, nothing Fantasti-specific
import ray
ray.init()  # connects to the Ray cluster on your Fantasti cluster

ds = ray.data.read_parquet("s3://acme-data/prompts/")
ds = ds.map_batches(  # GPU workers scale up from 0
    Embedder, concurrency=8, num_gpus=1, batch_size=256,
)
ds.write_parquet("s3://acme-data/embeddings/")
helm repo add kuberay https://ray-project.github.io/kuberay-helm/
helm install kuberay-operator kuberay/kuberay-operator

# ray-values.yaml: a GPU worker group, min 0
helm install raycluster kuberay/ray-cluster -f ray-values.yaml
Workers from zero. Illustrative.

Plain Ray code calls ray.init() and maps Embedder over a dataset with 8 workers of 1 GPU each. It connects to the Ray head that KubeRay runs on your Fantasti cluster. The GPU worker group starts at zero workers, scales to 8 while the job runs and returns to zero. Connections: Your code to Ray head; Ray head to GPU worker group.

  • DriverYour codePlain Ray
    ray.init() # plain Rayds.map_batches(Embedder, concurrency=8, num_gpus=1)
  • Ray headKubeRay
    kind
    RayCluster
    runs on
    your cluster
    workers
    min 0 · max 8
  • GPU worker groupFrom zero
    1 GPU each
  • DriverYour codePlain Ray
    ray.init() # plain Rayds.map_batches(Embedder, concurrency=8, num_gpus=1)
  • Ray headKubeRay
    kind
    RayCluster
    runs on
    your cluster
    workers
    min 0 · max 8
  • GPU worker groupFrom zero
    1 GPU each
  • DriverYour codePlain Ray
    ray.init() # plain Rayds.map_batches(Embedder, concurrency=8, num_gpus=1)
  • Ray headKubeRay
    kind
    RayCluster
    runs on
    your cluster
    workers
    min 0 · max 8
  • GPU worker groupFrom zero
    1 GPU each
  • Scales up
  • Your job
API One per cluster Reviewed 2026-10-09

1 Kubernetes API per cluster, and it is yours.

One isolated cluster per account
EQ 01 Source · Fantasti platform configuration
One cluster, every interface. Illustrative · example names.

5 tools reach one Kubernetes API, which is the control plane of your own cluster: Kubernetes (early access), Ray (early access), SkyPilot (early access), NVIDIA OSMO (planned), Applications (planned). Your kubeconfig points at it, and kubectl get pods -A lists your workloads and your cluster's system pods, nothing else. The API schedules onto dedicated GPU nodes that share one InfiniBand fabric, one of them from a spot pool with a max price. Every node reports runs to managed MLflow. Connections: Kubernetes to Kubernetes API: kubeconfig; Ray to Kubernetes API; SkyPilot to Kubernetes API: kube context; NVIDIA OSMO to Kubernetes API; Applications to Kubernetes API; Kubernetes API to gpu-node-01; Kubernetes API to gpu-node-02; Kubernetes API to gpu-node-03: spot pool, max price; gpu-node-01 to gpu-node-02; gpu-node-02 to gpu-node-03; One InfiniBand fabric to Managed MLflow.

One InfiniBand fabricDedicated GPU nodes
  • Kubernetes

    Early access

    $ kubectl get pods -A
    • Helm
    • operators
  • Ray

    Early access

    ray.init()
    • KubeRay
    • from zero
  • SkyPilot

    Early access

    infra: k8s/fantasti-acme
    • sky launch
    • sky serve
  • NVIDIA OSMO (planned)

    Planned

    $ osmo workflow submit
    • OSMO pools
  • Applications (planned)

    Planned

    fx.apps.install(…)
    • JupyterLab
    • vLLM
  • Kubernetes APIOne per cluster
    control plane
    yours, one per cluster
    kubeconfig
    fantasti-acme.yaml
    context
    fantasti-acme
    GPU nodes
    dedicated
    $ kubectl get pods -ANAMESPACE NAME STATUSdefault kuberay-operator Runningdefault raycluster-head Runningkube-system coredns Running
    API server
    1
    GPU nodes
    3
    GPUs
    24
  • gpu-node-01On-demand
    HGX H200 8-GPU baseboard, line drawing
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • gpu-node-02On-demand
    HGX H200 8-GPU baseboard, line drawing
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • gpu-node-03Spot
    HGX H200 8-GPU baseboard, line drawing
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • Managed MLflow
    tracking
    experiments, runs
    registry
    model versions
    MLFLOW_TRACKING_URI
    • Workspaces
    • Jobs
    • Batch runs

    Set on every node.

One InfiniBand fabricDedicated GPU nodes
  • Kubernetes

    Early access

    kubectl
  • Ray

    Early access

    ray.init()
  • SkyPilot

    Early access

    sky launch
  • NVIDIA OSMO (planned)

    Planned

    osmo
  • Applications (planned)

    Planned

    fx.apps
  • Kubernetes APIOne per cluster
    control plane
    yours, one per cluster
    kubeconfig
    fantasti-acme.yaml
    context
    fantasti-acme
    GPU nodes
    dedicated
    $ kubectl get pods -ANAMESPACE NAME STATUSdefault kuberay-operator Runningdefault raycluster-head Runningkube-system coredns Running
  • gpu-node-01On-demand
    HGX H200 8-GPU baseboard, line drawing
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • gpu-node-02Spot
    HGX H200 8-GPU baseboard, line drawing
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • Managed MLflow
    tracking
    experiments, runs
    registry
    model versions
    MLFLOW_TRACKING_URI
    • Workspaces
    • Jobs
    • Batch runs

    Set on every node.

One InfiniBand fabricDedicated GPU nodes
  • KubernetesEarly access
  • RayEarly access
  • SkyPilotEarly access
  • NVIDIA OSMOPlanned
  • ApplicationsPlanned
  • Kubernetes APIOne per cluster
    control plane
    yours, one per cluster
    kubeconfig
    fantasti-acme.yaml
    context
    fantasti-acme
    GPU nodes
    dedicated
    $ kubectl get pods -ANAMESPACE NAME STATUSdefault kuberay-operator Runningdefault raycluster-head Runningkube-system coredns Running
  • gpu-node-01On-demand
    8 GPUs
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • gpu-node-02Spot
    8 GPUs
    gpu.product
    NVIDIA-H200
    gpu.count
    8
  • Managed MLflowEarly access
    tracking
    experiments, runs
    registry
    model versions
    MLFLOW_TRACKING_URI
  • Your kubeconfig
  • Reports runs
  • Planned route
  • Planned
§03 SkyPilot Early access

Bring your SkyPilot YAML.

Fantasti issues a Kubernetes context for your cluster. Point SkyPilot at it and sky launch, sky jobs launch and sky serve work as they do anywhere else, with failover across the GPU types you list.

infra: k8s/fantasti-acme

SkyPilot’s own “workspaces” are team permissions. Fantasti Workspaces are GPU development environments; they are different things.

YAML
Preview API · subject to change
# sft.yaml: "fantasti-acme" is the kube context for your isolated cluster
name: sft-01
num_nodes: 2

resources:
  infra: k8s/fantasti-acme
  accelerators: {H200:8, B200:8}   # either works; SkyPilot takes what is free
  job_recovery:
    strategy: FAILOVER
    max_restarts_on_errors: 2

workdir: .
setup: pip install -r requirements.txt
run: |
  MASTER=$(echo "$SKYPILOT_NODE_IPS" | head -n1)
  torchrun --nnodes=$SKYPILOT_NUM_NODES --node_rank=$SKYPILOT_NODE_RANK \
    --nproc_per_node=$SKYPILOT_NUM_GPUS_PER_NODE \
    --master_addr=$MASTER --master_port=29500 train.py --resume
CLI
Preview API · subject to change
sky check kubernetes                      # finds your Fantasti context
sky gpus list --infra k8s/fantasti-acme   # GPUs in your cluster
sky jobs launch -n sft-01 sft.yaml     # recovers after node loss
§04 NVIDIA OSMO

NVIDIA OSMO for physical-AI workflows. Planned

Use OSMO on your Fantasti cluster to run simulation, synthetic-data and training workflows as one pipeline. Fantasti installs it and exposes your GPUs as OSMO pools.

YAML
Preview API · subject to change
# train.yaml: an OSMO workflow on your Fantasti cluster's GPU pool
workflow:
  name: policy-train
  resources:
    gpu_task:
      cpu: 16
      gpu: 1
      memory: 64Gi
      storage: 100Gi
  tasks:
  - name: train
    image: nvcr.io/nvidia/pytorch:24.01-py3
    command: ["bash", "-c", "python train.py --epochs 3"]
    resource: gpu_task
CLI
Preview API · subject to change
osmo pool list
osmo workflow submit train.yaml --pool fantasti-acme
One pipeline, one pool. Schematic · not to scale.

Three workflows, simulation, synthetic data and training, run in order as one pipeline. NVIDIA OSMO, installed by Fantasti on your cluster, submits their tasks to one pool named fantasti-acme, which is the GPU nodes of your cluster. The task train of the workflow policy-train holds 1 of the 8 GPUs of the first node. Connections: Simulation to Synthetic data; Synthetic data to Training; Simulation to NVIDIA OSMO; Synthetic data to NVIDIA OSMO; Training to NVIDIA OSMO; NVIDIA OSMO to gpu-node-01: submits tasks.

OSMO poolfantasti-acmeYour GPU nodes
  • 01Simulation
  • 02Synthetic data
  • 03Training
  • NVIDIA OSMOPlanned

    Installed by Fantasti on your cluster.

    workflow
    policy-train
    pool
    fantasti-acme
  • gpu-node-018 GPUs
    task train
  • gpu-node-028 GPUs
    free
OSMO poolfantasti-acme
  • 01Simulation
  • 02Synthetic data
  • 03Training
  • NVIDIA OSMOPlanned

    Installed by Fantasti on your cluster.

    workflow
    policy-train
    pool
    fantasti-acme
  • gpu-node-018 GPUs
    task train
  • gpu-node-028 GPUs
    free
  • One pipeline
  • Planned
§05 MLflow Early access

Managed MLflow.

Managed MLflow tracking and model registry, wired into Workspaces, Flows, Batch Inference, Serverless jobs and Clusters.

One run, on the record. Illustrative.

A workload on any of your nodes, in Workspaces, jobs or batch runs, already has MLFLOW_TRACKING_URI and credentials set. It logs the parameter lr and the metric eval_loss to the experiment sft-01, which holds 3 runs. Run 03 is registered as version 3 of the model sft-01. Connections: Any workload to sft-01: logs; sft-01 to sft-01: registers.

  • Any workloadOn your nodes
    • Workspaces
    • Jobs
    • Batch runs
    MLFLOW_TRACKING_URI
    set
    credentials
    set
    mlflow.log_params({"lr": 2e-5})mlflow.log_metric("eval_loss", …)
  • Experimentsft-013 runs
    run lr eval_loss03 2e-5 0.41202 5e-5 0.44701 1e-4 0.503
    artifacts
    checkpoints/last
  • Modelsft-01Registry
    v3 run 03v2 run 02v1 run 01
  • Any workloadOn your nodes
    • Workspaces
    • Jobs
    • Batch runs
    MLFLOW_TRACKING_URI
    set
    credentials
    set
    mlflow.log_params({"lr": 2e-5})mlflow.log_metric("eval_loss", …)
  • Experimentsft-013 runs
    run lr eval_loss03 2e-5 0.41202 5e-5 0.44701 1e-4 0.503
    artifacts
    checkpoints/last
  • Modelsft-01Registry
    v3 run 03v2 run 02v1 run 01
  • Any workloadOn your nodes
    • Workspaces
    • Jobs
    • Batch runs
    MLFLOW_TRACKING_URI
    set
    credentials
    set
    mlflow.log_params({"lr": 2e-5})mlflow.log_metric("eval_loss", …)
  • Experimentsft-013 runs
    run lr eval_loss03 2e-5 0.41202 5e-5 0.44701 1e-4 0.503
    artifacts
    checkpoints/last
  • Modelsft-01Registry
    v3 run 03v2 run 02v1 run 01
  • Registered
  • Logged by your code
Python
Preview API · subject to change
# MLFLOW_TRACKING_URI and credentials are already set
import mlflow

mlflow.set_experiment("sft-01")
with mlflow.start_run():
    mlflow.log_params({"lr": 2e-5, "epochs": 3})
    mlflow.log_metric("eval_loss", eval_loss)
    mlflow.log_artifact("checkpoints/last/config.json")
  • 01Tracking

    Experiments and runs

    Log parameters, metrics and artifacts from any workload.

  • 02Registry

    Model versions

    Register, version and promote models.

  • 03Wired in

    Set on every node

    Workspaces, jobs and batch runs get MLFLOW_TRACKING_URI and credentials.

§06 Applications 5 categories

Applications, one click from your GPUs. Planned

TAB 02Applications by category Source · Fantasti
Applications by category: 5 categories, 11 applications, and how each category runs.
Category Applications Managed app On your cluster Runs as
Notebooks and IDEs
  • JupyterLab
  • code-server
YesYes Managed app or on your cluster
Inference servers
  • vLLM
  • SGLang
  • NVIDIA Triton
NoYes On your cluster
Vector databases
  • Qdrant
YesYes Managed app or on your cluster
Interfaces
  • Open WebUI
  • ComfyUI
YesNo Managed app
Schedulers
  • Ray Cluster
  • SkyPilot API server
  • NVIDIA OSMO
NoYes On your cluster
Python
Preview API · subject to change
app = fx.apps.install("jupyterlab", cluster="acme-train", gpu="L40S")
print(app.url)
§07 Q&A 5 questions

Questions and answers.

Do I need to change my Ray or SkyPilot code?

No. Your cluster exposes a standard Kubernetes API, and Ray and SkyPilot use it the way their docs describe.

Is my cluster shared with other accounts?

No. It has its own control plane and dedicated GPU nodes.

Can my scheduler run on spot capacity?

Yes. Add spot node pools with a max price to your cluster. A SkyPilot task that lists more than one GPU type fails over across them. How spot works

Which integrations are in early access?

Kubernetes, Ray, SkyPilot and MLflow are in early access. NVIDIA OSMO and Applications are planned.

What about Metaflow?

Metaflow runs as Fantasti Flows, which is planned. Flows
Sheet
01
Title
Integrations · applications and answers
Reviewed
2026-10-09
§08 Request Early access

Keep your scheduler. Change where it runs.

Integrations is in early access. Request access and tell us which scheduler you run.