fx.jobs.create(image=…, gpu="H200:8")RunServerless Early access
Run a container on a GPU.
Skip the cluster.
Serverless runs your container as a job that finishes or an endpoint that serves requests. The GPU is provisioned for it and metered per second, and compute billing stops when the job ends or the endpoint stops.
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
job = fx.jobs.create(
image="ghcr.io/acme/train:latest",
command=["python", "train.py", "--epochs", "3"],
gpu="H200:8",
timeout="24h",
capacity=fantasti.Spot(max_price="follow"),
mlflow="support-sft", # experiment name
)
for line in job.logs(follow=True):
print(line) Output · illustrative
job_4n7w placing H200:8 · spot · max_price follow
job_4n7w running mlflow experiment support-sft
epoch 1/3 · checkpoint saved
epoch 2/3 · checkpoint saved
epoch 3/3 · checkpoint saved
job_4n7w exited 0 container disk removed · metering stopped $ fantasti job submit --image ghcr.io/acme/train:latest \
--gpu H200:8 --timeout 24h --spot --max-price follow \
--mlflow support-sft -- python train.py --epochs 3
$ fantasti job logs job_4n7w --follow Output · illustrative
job_4n7w placing H200:8 · spot · max_price follow
job_4n7w running mlflow experiment support-sft
epoch 1/3 · checkpoint saved
epoch 2/3 · checkpoint saved
epoch 3/3 · checkpoint saved
job_4n7w exited 0 container disk removed · metering stopped {
"mcpServers": {
"fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
}
} Job tools · early access
jobs_create confirms a launch above the key's per-call budget
jobs_logs curl -X POST https://api.fantasti.ai/v1/jobs \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"image": "ghcr.io/acme/train:latest",
"command": ["python", "train.py", "--epochs", "3"],
"gpu": "H200:8",
"timeout": "24h",
"capacity": { "mode": "spot", "max_price": "follow" },
"mlflow": "support-sft"
}'
# { "id": "job_4n7w…", "status": "placing" } - Shapes
- job · endpoint
- Capacity
- on-demand · spot with max_price
- Metering
- per second, while it runs
- Tracking
- managed MLflow
Two shapes.
For engineers who want a GPU for a task, not a machine.
- Job Runs to completion Training, fine-tuning, preprocessing, evaluation. Restarts on failure.
- Endpoint Serves requests at a URL A model or app behind HTTPS, with a token if you want one. Stop it and compute billing stops.
fx.endpoints.create(image=…, port=8000) - Dev Interactive work Use a Workspace. Same image, same bill.
/workspaces
Jobs: start, run, finish.
Give Fantasti an image, a command and a GPU. The job runs until it exits or reaches its timeout, restarts if it fails, and its container disk is removed when it ends. Jobs can run on spot with a max price.
- Timeoutsset a limit per job, up to 7 days
- Spotfollow the market or set a max price
- Trackedlog to managed MLflow
Your agent can do this. jobs_create and jobs_logs are MCP tools.
MCP server Early access
The record of the job job_4n7w: image ghcr.io/acme/train:latest, 8 H200 GPUs, spot capacity with max_price follow, a timeout of 24h and the MLflow experiment support-sft. The job is running. It is metered per second from its start until its command exits or its timeout passes, whichever comes sooner, and it restarts if it fails. Its container disk is removed when it ends. Connections: job_4n7w… to Running: placed; Running to Container disk: writes; Running to Managed MLflow: logs.
- image
- ghcr.io/acme/train:latest
- gpu
- H200:8
- capacity
- spot · max_price follow
- timeout
- 24h
- mlflow
- experiment support-sft
- state
- running
placing H200:8 · spotrunning mlflow experiment support-sftepoch 1/3 · checkpoint saved- restarts on failure
- timeout 24h
HGX H200 · 8 GPUs
Lives as longas the job
Experiment
support-sft
- image
- ghcr.io/acme/train:latest
- gpu
- H200:8
- capacity
- spot · max_price follow
- timeout
- 24h
- mlflow
- experiment support-sft
- state
- running
placing H200:8 · spotrunning mlflow experiment support-sftepoch 1/3 · checkpoint savedHGX H200 · 8 GPUs
- GPUs
- 8
- Timeout
- 24h
- restarts on failure
- timeout 24h
Lives as long as the jobRemoved at exit
Experiment
support-sft
- image
- ghcr.io/acme/train:latest
- gpu
- H200:8
- capacity
- spot · max_price follow
- timeout
- 24h
- mlflow
- experiment support-sft
- state
- running
placing H200:8 · spotrunning mlflow experiment support-sftepoch 1/3 · checkpoint saved- restarts on failure
- timeout 24h
HGX H200 · 8 GPUs
- to
- Managed MLflow
Lives as longas the job
Experiment
support-sft
- This job
7 days is the longest timeout a job can set.
timeout="24h"The job on this page
A job ends when its command exits or at its timeout, whichever comes sooner. Its container disk is removed, and metering stops with it.
Endpoints: a model at a URL.
Serve a container at a managed HTTPS address, protected by a token. Stop it when you are not using it and compute billing stops with it. During early access each endpoint runs as one replica.
ep = fx.endpoints.create(
image="ghcr.io/acme/serve:latest",
args=["--model", "s3://acme-models/support-sft"],
gpu="H100:1",
port=8000,
auth="token",
)
print(ep.url) # https://ep-… (managed HTTPS)
ep.stop() # compute billing stops; start it again later Output · illustrative
ep_9c3t placing H100:1 · on-demand
ep_9c3t running https://ep-… · token required · 1 replica
ep_9c3t stopped compute billing stopped $ fantasti endpoint create support-sft \
--image ghcr.io/acme/serve:latest \
--gpu H100:1 --port 8000 --auth token \
-- --model s3://acme-models/support-sft
$ fantasti endpoint stop support-sft # start it again later Output · illustrative
ep_9c3t placing H100:1 · on-demand
ep_9c3t running https://ep-… · token required · 1 replica
ep_9c3t stopped compute billing stopped curl -X POST https://api.fantasti.ai/v1/endpoints \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"image": "ghcr.io/acme/serve:latest",
"args": ["--model", "s3://acme-models/support-sft"],
"gpu": "H100:1",
"port": 8000,
"auth": "token"
}'
# { "id": "ep_9c3t…", "status": "placing" } - Addressmanaged HTTPS, read from
ep.url - Token
auth="token"protects the address - Stoppeda stopped endpoint costs nothing
A client sends a request and the token over HTTPS to the endpoint's managed address. The address checks the token and passes the request to your container, which listens on port 8000 on one H100:1 replica and runs until you stop it. Connections: Your client to Managed address: request; Managed address to Your container: 1 replica.
Any HTTPclient
- sends
- request + token
Checksthe token
- url
- https://ep-…
- auth
- "token"
Runs untilyou stop it
- port
- 8000
- replicas
- 1
- sends
- request
- + token
Any HTTP client
- url
- https://ep-…
- auth
- "token"
Checks the token
- gpu
- H100:1
- port
- 8000
- replicas
- 1
Runs untilyou stop it
Any HTTP client
- sends
- request + token
Checksthe token
- url
- https://ep-…
- auth
- "token"
Runs untilyou stop it
- port
- 8000
- replicas
- 1
Over time, the endpoint ep_9c3t is placed and then runs, metered per second. When you stop it, nothing is metered. When you start it again it runs and is metered per second, until you stop it.
Spot for jobs, with your max price.
Every GPU type has a spot price that moves with supply and demand and can change as often as every 15 minutes. Pick one of two modes for each request.
-
Mode 01
Follow the market
fantasti.Spot(max_price="follow")No ceiling. You pay the spot price for each interval, and Fantasti never stops your work because the price rose. Spot capacity can still be reclaimed.
-
Mode 02Drawn in the plot
Set a max price
fantasti.Spot(max_price=2.40)Choose the most you will pay per GPU-hour. Your nodes run while the spot price is at or below it, and you pay the spot price, not your max. If the spot price rises above your max, those nodes stop.
- Job
- H100:1
- Max price
- $2.40 /GPU·hr
- Ran
- 51 h 15 m of 72 h
- Stopped
- 4×
- Resumed
- 4×
Blocked at start
Illustrative series, not Fantasti market data. A step line shows a spot price for one GPU type over 72 hours in 15-minute steps, between $1.91 and $2.72 per GPU-hour. With the max price at $2.40, the nodes run 51 h 15 m of 72 hours, stop 4 times when the price rises above the max, and resume 4 times when it falls back. You pay the spot price for each interval, not the max.
Limits in early access
- Job timeout
- Up to 7 days
- Set per job. A job ends at its exit or at its timeout.
- Endpoint replicas
- One per endpoint
- One replica serves every request during early access.
- Max price
- Per request, per GPU-hour
- It limits what you pay. It does not reserve capacity.
Metered per second, at the rate of the GPU.
A job is metered from its start to its exit, an endpoint while it runs. Both appear as lines on your Fantasti statement, at the rate of the GPU they run on.
| Shape | Runs on | Capacity | Per GPU-hour | Per GPU-second | Metered |
|---|---|---|---|---|---|
| Job job_4n7w | 8 × H200 SXM | Spot max_price="follow" | From $0.87 | From $0.000242 | From its start to its exit |
| Endpoint ep_9c3t | 1 × H100 SXM5 | On-demand | $5.40 | $0.001500 | While it runs. Nothing while it is stopped. |
Metered by the second, billed hourly. "From" is the lowest spot price for that GPU. The spot price moves with supply and demand. The spot price always stays below the on-demand rate for the same GPU. All prices
Serverless, Batch Inference or Compute?
| Question | Serverless Early access | Batch Inference | GPU Instances | GPU Clusters |
|---|---|---|---|---|
| You bring | A container | A model and a dataset | An image, or your own stack | A scheduler, or nothing but SSH |
| Fantasti runs | One job or one endpoint | A sharded inference run | One machine with 1 or 8 GPUs | Many 8-GPU nodes on one fabric |
| Ends | When the job exits or you stop the endpoint | When every shard is done | When you stop it or its term ends | When you stop it or its term ends |
| Keeps | Nothing on the container disk | Outputs and checkpoints | Its volumes | Its volumes and shared filesystem |
| Capacity | On-demand or spot | Spot or on-demand | On-demand, spot, reserved | On-demand, spot workers, reserved |
| Use it when | The job is one container | The job is one model over one dataset | You need a GPU and root | The job spans nodes |
Questions about Serverless.
Do I manage any VMs?
How is it billed?
Can a job run on spot?
Do endpoints autoscale?
- Sheet
- 01 / 01
- Title
- Serverless: rates, fit and answers
- Metering
- Per second
- Reviewed
- 2026-10-10
Bring a container.
We bring the GPU.
Serverless is in early access. Companies can request a place in the first cohort. Tell us what you want to run.