RunServerless Early access

Run a container on a GPU.
Skip the cluster.

Serverless runs your container as a job that finishes or an endpoint that serves requests. The GPU is provisioned for it and metered per second, and compute billing stops when the job ends or the endpoint stops.

Being built for the first cohort. Join by request.
Preview API · subject to change
import fantasti

fx = fantasti.Client()        # reads FANTASTI_API_KEY

job = fx.jobs.create(
    image="ghcr.io/acme/train:latest",
    command=["python", "train.py", "--epochs", "3"],
    gpu="H200:8",
    timeout="24h",
    capacity=fantasti.Spot(max_price="follow"),
    mlflow="support-sft",     # experiment name
)
for line in job.logs(follow=True):
    print(line)

Output · illustrative

job_4n7w  placing   H200:8 · spot · max_price follow
job_4n7w  running   mlflow experiment support-sft
epoch 1/3 · checkpoint saved
epoch 2/3 · checkpoint saved
epoch 3/3 · checkpoint saved
job_4n7w  exited 0  container disk removed · metering stopped
$ fantasti job submit --image ghcr.io/acme/train:latest \
    --gpu H200:8 --timeout 24h --spot --max-price follow \
    --mlflow support-sft -- python train.py --epochs 3
$ fantasti job logs job_4n7w --follow

Output · illustrative

job_4n7w  placing   H200:8 · spot · max_price follow
job_4n7w  running   mlflow experiment support-sft
epoch 1/3 · checkpoint saved
epoch 2/3 · checkpoint saved
epoch 3/3 · checkpoint saved
job_4n7w  exited 0  container disk removed · metering stopped
{
  "mcpServers": {
    "fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
  }
}

Job tools · early access

jobs_create   confirms a launch above the key's per-call budget
jobs_logs
curl -X POST https://api.fantasti.ai/v1/jobs \
  -H "Authorization: Bearer $FANTASTI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "image": "ghcr.io/acme/train:latest",
        "command": ["python", "train.py", "--epochs", "3"],
        "gpu": "H200:8",
        "timeout": "24h",
        "capacity": { "mode": "spot", "max_price": "follow" },
        "mlflow": "support-sft"
      }'
# { "id": "job_4n7w…", "status": "placing" }
Shapes
job · endpoint
Capacity
on-demand · spot with max_price
Metering
per second, while it runs
Tracking
managed MLflow
§01 Two shapes Job · Endpoint

Two shapes.

For engineers who want a GPU for a task, not a machine.

  • Job Runs to completion Training, fine-tuning, preprocessing, evaluation. Restarts on failure. fx.jobs.create(image=…, gpu="H200:8")
  • Endpoint Serves requests at a URL A model or app behind HTTPS, with a token if you want one. Stop it and compute billing stops. fx.endpoints.create(image=…, port=8000)
  • Dev Interactive work Use a Workspace. Same image, same bill. /workspaces
§02 Jobs Illustrative

Jobs: start, run, finish.

Give Fantasti an image, a command and a GPU. The job runs until it exits or reaches its timeout, restarts if it fails, and its container disk is removed when it ends. Jobs can run on spot with a max price.

  • Timeoutsset a limit per job, up to 7 days
  • Spotfollow the market or set a max price
  • Trackedlog to managed MLflow

Your agent can do this. jobs_create and jobs_logs are MCP tools.

One job, from request to exit. Illustrative.

The record of the job job_4n7w: image ghcr.io/acme/train:latest, 8 H200 GPUs, spot capacity with max_price follow, a timeout of 24h and the MLflow experiment support-sft. The job is running. It is metered per second from its start until its command exits or its timeout passes, whichever comes sooner, and it restarts if it fails. Its container disk is removed when it ends. Connections: job_4n7w… to Running: placed; Running to Container disk: writes; Running to Managed MLflow: logs.

  • Recordjob_4n7w…console · API
    image
    ghcr.io/acme/train:latest
    gpu
    H200:8
    capacity
    spot · max_price follow
    timeout
    24h
    mlflow
    experiment support-sft
    state
    running
  • JobRunningH200:8 · spot
    placing H200:8 · spotrunning mlflow experiment support-sft epoch 1/3 · checkpoint saved
    • restarts on failure
    • timeout 24h
    HGX H200 8-GPU baseboard, line drawingHGX H200 · 8 GPUs
  • Container diskremoved at exit

    Lives as longas the job

  • Managed MLflowtracked

    Experiment

    support-sft
  • Recordjob_4n7w…console · API
    image
    ghcr.io/acme/train:latest
    gpu
    H200:8
    capacity
    spot · max_price follow
    timeout
    24h
    mlflow
    experiment support-sft
    state
    running
  • JobRunningH200:8 · spot
    placing H200:8 · spotrunning mlflow experiment support-sft epoch 1/3 · checkpoint saved
    HGX H200 8-GPU baseboard, line drawingHGX H200 · 8 GPUs
    GPUs
    8
    Timeout
    24h
    • restarts on failure
    • timeout 24h
  • Container disk

    Lives as long as the jobRemoved at exit

  • Managed MLflow

    Experiment

    support-sft
  • Recordjob_4n7w…console · API
    image
    ghcr.io/acme/train:latest
    gpu
    H200:8
    capacity
    spot · max_price follow
    timeout
    24h
    mlflow
    experiment support-sft
    state
    running
  • JobRunningH200:8 · spot
    placing H200:8 · spotrunning mlflow experiment support-sft epoch 1/3 · checkpoint saved
    • restarts on failure
    • timeout 24h
    HGX H200 8-GPU baseboard, line drawingHGX H200 · 8 GPUs
    to
    Managed MLflow
  • Container diskremoved at exit

    Lives as longas the job

  • Managed MLflowtracked

    Experiment

    support-sft
  • This job

7 days is the longest timeout a job can set.

Timeout · set per job
EQ 01 Source · Serverless API preview

timeout="24h"The job on this page

A job ends when its command exits or at its timeout, whichever comes sooner. Its container disk is removed, and metering stops with it.

§03 Endpoints Preview API

Endpoints: a model at a URL.

Serve a container at a managed HTTPS address, protected by a token. Stop it when you are not using it and compute billing stops with it. During early access each endpoint runs as one replica.

Preview API · subject to change
ep = fx.endpoints.create(
    image="ghcr.io/acme/serve:latest",
    args=["--model", "s3://acme-models/support-sft"],
    gpu="H100:1",
    port=8000,
    auth="token",
)
print(ep.url)    # https://ep-…  (managed HTTPS)
ep.stop()        # compute billing stops; start it again later

Output · illustrative

ep_9c3t  placing   H100:1 · on-demand
ep_9c3t  running   https://ep-… · token required · 1 replica
ep_9c3t  stopped   compute billing stopped
$ fantasti endpoint create support-sft \
    --image ghcr.io/acme/serve:latest \
    --gpu H100:1 --port 8000 --auth token \
    -- --model s3://acme-models/support-sft
$ fantasti endpoint stop support-sft    # start it again later

Output · illustrative

ep_9c3t  placing   H100:1 · on-demand
ep_9c3t  running   https://ep-… · token required · 1 replica
ep_9c3t  stopped   compute billing stopped
curl -X POST https://api.fantasti.ai/v1/endpoints \
  -H "Authorization: Bearer $FANTASTI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "image": "ghcr.io/acme/serve:latest",
        "args": ["--model", "s3://acme-models/support-sft"],
        "gpu": "H100:1",
        "port": 8000,
        "auth": "token"
      }'
# { "id": "ep_9c3t…", "status": "placing" }
  • Addressmanaged HTTPS, read from ep.url
  • Tokenauth="token" protects the address
  • Stoppeda stopped endpoint costs nothing
A request to an endpoint. Schematic · not to scale.

A client sends a request and the token over HTTPS to the endpoint's managed address. The address checks the token and passes the request to your container, which listens on port 8000 on one H100:1 replica and runs until you stop it. Connections: Your client to Managed address: request; Managed address to Your container: 1 replica.

  • Your client

    Any HTTPclient

    sends
    request + token
  • Managed addressHTTPS

    Checksthe token

    url
    https://ep-…
    auth
    "token"
  • Your containerH100:1

    Runs untilyou stop it

    port
    8000
    replicas
    1
  • Your client
    sends
    request
    + token

    Any HTTP client

  • Managed address
    url
    https://ep-…
    auth
    "token"

    Checks the token

  • Your container (highlighted)
    gpu
    H100:1
    port
    8000
    replicas
    1

    Runs untilyou stop it

  • Your client

    Any HTTP client

    sends
    request + token
  • Managed addressHTTPS

    Checksthe token

    url
    https://ep-…
    auth
    "token"
  • Your containerH100:1

    Runs untilyou stop it

    port
    8000
    replicas
    1

Over time, the endpoint ep_9c3t is placed and then runs, metered per second. When you stop it, nothing is metered. When you start it again it runs and is metered per second, until you stop it.

§04 Spot and limits Illustrative

Spot for jobs, with your max price.

Every GPU type has a spot price that moves with supply and demand and can change as often as every 15 minutes. Pick one of two modes for each request.

  1. Mode 01

    Follow the market

    fantasti.Spot(max_price="follow")

    No ceiling. You pay the spot price for each interval, and Fantasti never stops your work because the price rose. Spot capacity can still be reclaimed.

  2. Mode 02Drawn in the plot

    Set a max price

    fantasti.Spot(max_price=2.40)

    Choose the most you will pay per GPU-hour. Your nodes run while the spot price is at or below it, and you pay the spot price, not your max. If the spot price rises above your max, those nodes stop.

PLT 01Spot price and a max price over 72 hours Illustrative series · not Fantasti market data
Job
H100:1
Max price
$2.40 /GPU·hr
Ran
51 h 15 m of 72 h
Stopped
4×
Resumed
4×

Illustrative series, not Fantasti market data. A step line shows a spot price for one GPU type over 72 hours in 15-minute steps, between $1.91 and $2.72 per GPU-hour. With the max price at $2.40, the nodes run 51 h 15 m of 72 hours, stop 4 times when the price rises above the max, and resume 4 times when it falls back. You pay the spot price for each interval, not the max.

Limits in early access

Job timeout
Up to 7 days
Set per job. A job ends at its exit or at its timeout.
Endpoint replicas
One per endpoint
One replica serves every request during early access.
Max price
Per request, per GPU-hour
It limits what you pay. It does not reserve capacity.
§05 Rates Prices reviewed 2026-10-10

Metered per second, at the rate of the GPU.

A job is metered from its start to its exit, an endpoint while it runs. Both appear as lines on your Fantasti statement, at the rate of the GPU they run on.

TAB 01Rates for the job and the endpoint on this page Source · Fantasti price list Reviewed 2026-10-10
Rates for the job and the endpoint on this page: what each runs on, its capacity mode, the price per GPU-hour and per GPU-second, and when it is metered
Shape Runs on Capacity Per GPU-hour Per GPU-second Metered
Job job_4n7w 8 × H200 SXM Spot max_price="follow" From $0.87 From $0.000242 From its start to its exit
Endpoint ep_9c3t 1 × H100 SXM5 On-demand $5.40 $0.001500 While it runs. Nothing while it is stopped.

Metered by the second, billed hourly. "From" is the lowest spot price for that GPU. The spot price moves with supply and demand. The spot price always stays below the on-demand rate for the same GPU. All prices

§06 Which product 4 of 7 products

Serverless, Batch Inference or Compute?

TAB 02Serverless, Batch Inference or Compute Source · Fantasti
Serverless, Batch Inference or Compute
Question Serverless Early access Batch Inference GPU Instances GPU Clusters
You bring A container A model and a dataset An image, or your own stack A scheduler, or nothing but SSH
Fantasti runs One job or one endpoint A sharded inference run One machine with 1 or 8 GPUs Many 8-GPU nodes on one fabric
Ends When the job exits or you stop the endpoint When every shard is done When you stop it or its term ends When you stop it or its term ends
Keeps Nothing on the container disk Outputs and checkpoints Its volumes Its volumes and shared filesystem
Capacity On-demand or spot Spot or on-demand On-demand, spot, reserved On-demand, spot workers, reserved
Use it when The job is one container The job is one model over one dataset You need a GPU and root The job spans nodes
§07 Q&A 4 questions

Questions about Serverless.

Do I manage any VMs?

No. You provide a container; Fantasti runs it on a GPU and removes it when it is done.

How is it billed?

Per second while a job or endpoint runs, on your Fantasti statement. A stopped endpoint costs nothing.

Can a job run on spot?

Yes. Follow the market or set a max price. If a spot job is stopped, it restarts when capacity at your price returns.

Do endpoints autoscale?

Not during early access. Each endpoint runs one replica.
Sheet
01 / 01
Title
Serverless: rates, fit and answers
Metering
Per second
Reviewed
2026-10-10
§08 Request Early access

Bring a container.
We bring the GPU.

Serverless is in early access. Companies can request a place in the first cohort. Tell us what you want to run.