RunBatch Inference Private Preview
Stopped,
not lost.
Run async inference, evaluation and data-processing jobs on spot capacity. Workers checkpoint as they go, retry with backoff, and resume on new capacity when a node is reclaimed.
fx.batch.create(input="s3://…") The run is queued, placed and runs on spot capacity with a max price. The spot price rises above the max: the nodes get a stop notice and stop, and their disks are kept. When the price is back at or below the max the run resumes, restores checkpoint 0412, runs again and completes. Inputs, checkpoints and results stay in your own bucket, s3://acme-evals/: workers read prompts/*.jsonl, write a checkpoint to ckpt/ every 5m, restore checkpoint 0412 from it after the stop, and write results to results/. Connections: Your request to Queue; Queue to bat_2kq8x1: placed; prompts/ to bat_2kq8x1: read; bat_2kq8x1 to ckpt/: write; ckpt/ to bat_2kq8x1: restore; bat_2kq8x1 to results/: results.
s3://acme-evals/Your bucketfx.batch.create(gpu="H100:8",capacity=fantasti.Spot(max_price=2.40),checkpoint={"every": "5m"},priority="low")- order
- 06 of 06
- waits for
- spot ≤ 2.40
queued priority="low"placed us-central · fabric f · spotrunning ckpt 0001…0411stopped price above max · stop notice · disks keptresuming price back at or below maxrestored checkpoint 0412running ckpt 0413…0626HGX H100 · 8 GPUs
*.jsonl- restored
- 0412
done
s3://acme-evals/Your bucketfx.batch.create(gpu="H100:8",capacity=fantasti.Spot(max_price=2.40),checkpoint={"every": "5m"},priority="low")- order
- 06 of 06
- max price
- 2.40
queuedplaced us-central · fabric f · spotrunning ckpt 0001…0411stopped price above maxresuming price back at or below maxrestored checkpoint 0412running ckpt 0413…0626HGX H100 · 8 GPUs
*.jsonl- restored
- 0412
done
fx.batch.create(gpu="H100:8",capacity=fantasti.Spot(max_price=2.40),checkpoint={"every": "5m"},priority="low")- order
- 06 of 06
- waits for
- spot ≤ 2.40
queued priority="low"placed us-central · fabric f · spotrunning ckpt 0001…0411stopped price above maxresuming price back at or below maxrestored checkpoint 0412running ckpt 0413…0626HGX H100 · 8 GPUs
- to
- results/
- written
- every 5m
- restored
- 0412
done
- Restore from the checkpoint
- This run
- Reads and writes
Mechanics
checkpoint={"every": "5m"}- Checkpointing
- Workers write progress at an interval you set and resume from the last checkpoint after a stop.
retries={"max": 3, …}- Backoff retries
- Failed or stopped work retries with exponential backoff, up to the limit you set.
stopped → resuming → running- Workers that survive a stop
- A worker that is reclaimed loses at most the work since its last checkpoint.
priority="low"- Queue priorities
- Urgent jobs run first. Low-priority jobs wait for capacity under your price ceiling.
Submit a job.
Collect results
in your bucket.
Inputs and outputs live in your object storage. You pass URIs; workers read and write there.
- TrackedRuns log to managed MLflow. Early access
- From a flowA flow step can start a batch run and wait for it. Planned
- From your agentBatch tools on the MCP server. Planned
import fantasti
fx = fantasti.Client() # reads FANTASTI_API_KEY
run = fx.batch.create(
image="ghcr.io/acme/llm-eval:2026.10", # your container
input="s3://acme-evals/prompts/*.jsonl",
output="s3://acme-evals/results/",
gpu="H100:8",
capacity=fantasti.Spot(max_price=2.40), # USD per GPU-hour
checkpoint={"every": "5m"}, # the resume point
retries={"max": 3, "backoff": "exponential"},
priority="low", # high, normal, low
)
print(run.id, run.status) # bat_2kq8x1 queued $ fantasti batch create \
--image ghcr.io/acme/llm-eval:2026.10 \
--input "s3://acme-evals/prompts/*.jsonl" \
--output s3://acme-evals/results/ \
--gpu H100:8 --spot --max-price 2.40 \
--checkpoint 5m --retries 3 --priority low
$ fantasti batch get bat_2kq8x1 --watch Output · illustrative
queued priority="low"
placed us-central · fabric f · spot
running ckpt 0001…0411
stopped price above max · stop notice · disks kept
resuming price back at or below max
restored checkpoint 0412 · s3://acme-evals/ckpt/
running ckpt 0413…0626
done s3://acme-evals/results/ // .mcp.json · project scope, OAuth sign-in
{
"mcpServers": {
"fantasti": {
"type": "http",
"url": "https://mcp.fantasti.ai/mcp"
}
}
} Batch tools · planned
batch_create
batch_get
batch_cancel asks for confirmation curl -X POST https://api.fantasti.ai/v1/batch \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Idempotency-Key: evals-1009" \
-H "Content-Type: application/json" \
-d '{
"image": "ghcr.io/acme/llm-eval:2026.10",
"input": "s3://acme-evals/prompts/*.jsonl",
"output": "s3://acme-evals/results/",
"gpu": "H100:8",
"capacity": { "mode": "spot", "max_price": "2.40",
"currency": "USD", "unit": "gpu_hour" },
"checkpoint": { "every": "5m" },
"retries": { "max": 3, "backoff": "exponential" },
"priority": "low"
}'
# { "id": "bat_2kq8x1", "status": "queued" } Set a ceiling.
Run while the market is under it.
Drag the max price line, click the plot, or use the arrow keys.
- Run
- H100:8
- Max price
- $2.40 /GPU·hr
- Ran
- 51 h 15 m of 72 h
- Stopped
- 4×
- Resumed
- 4×
Blocked at start
Illustrative series, not Fantasti market data. A step line shows a spot price for one GPU type over 72 hours in 15-minute steps, between $1.91 and $2.72 per GPU-hour. With the max price at $2.40, the nodes run 51 h 15 m of 72 hours, stop 4 times when the price rises above the max, and resume 4 times when it falls back. You pay the spot price for each interval, not the max.
1 checkpoint interval is the most a reclaimed worker loses.
In the illustrative run at the top of this page the nodes stop once. The run restores checkpoint 0412 and carries on from there. Nothing before it runs again.
Illustrative. A worker writes a checkpoint at every interval (every 5m in the sample on this page). Its nodes stop partway through the interval after checkpoint 0412. After the run resumes, the worker restores that checkpoint and repeats only the work since it, which is at most one interval. Everything before the last checkpoint is kept.
- Kept
- Everything up to checkpoint 0412
- Repeated
- At most one interval (5m in this sample)
- Compute charged while stopped
- None. Storage continues.
Queue priorities.
| Priority | Scheduling | Use for |
|---|---|---|
| high | Scheduled first, on any capacity your policy allows | Release-blocking evals |
| normal | Scheduled after high-priority work, on spot capacity by default | Daily inference and scoring |
| low | Waits for spot capacity under your price ceiling | Backfills, re-scoring, data processing |
Illustrative. 6 runs from one organisation, in scheduling order, against a quota of 48 H100 GPUs. release-evals (high, 16 GPUs), hotfix-evals (high, 8 GPUs), daily-scoring (normal, 24 GPUs) fit inside the quota and are running. embed-docs (normal, 16 GPUs) waits for quota; backfill-q3 (low, 16 GPUs) waits for spot under max; rescore-evals (low, 8 GPUs) waits for spot under max.
| Order | Run | Priority | Capacity | GPUs | State |
|---|---|---|---|---|---|
| 01 | release-evals (bat_9c1r4m) | high | on-demand | 16 | running |
| 02 | hotfix-evals (bat_7m3d0k) | high | spot ≤ 2.40 | 8 | running |
| 03 | daily-scoring (bat_5t0w2p) | normal | spot ≤ 2.40 | 24 | running |
| 04 | embed-docs (bat_3f6q8a) | normal | spot ≤ 2.40 | 16 | queued, waits for quota |
| 05 | backfill-q3 (bat_1x8n5c) | low | spot ≤ 2.40 | 16 | queued, waits for spot under max |
| 06 | rescore-evals (bat_2kq8x1) | low | spot ≤ 2.40 | 8 | queued, waits for spot under max |
One organisation, one quota of 48 H100 GPUs. Each bar is the GPUs a run asks for, placed after the runs scheduled ahead of it. Bars inside the quota are running; bars past it are queued.
Pay for worker seconds.
Workers are metered by the second at the rate of the capacity they run on: the on-demand list price, or the spot price in effect for each interval. Time a worker spends stopped is not running and is not billed.
- MeterMetered by the second, billed hourly.
- On demand$5.40 per GPU-hour for H100 SXM5, the list price.
- SpotFrom $0.87 per GPU-hour for H100 SXM5. You pay the spot price of each interval, not your max.
A stopped worker keeps its disks, so its storage continues.
| Granularity | Seconds charged | Formula | Charge |
|---|---|---|---|
| Per hour | 3,600 s | 3,600 s × $5.40 / 3,600 s = $5.40 | $5.40 |
| Per minute | 300 s | 300 s × $5.40 / 3,600 s = $0.45 | $0.45 |
| Per second | 250 s | 250 s × $5.40 / 3,600 s = $0.38 | $0.38 |
Fantasti meters by the second and bills hourly (Terms §3). A 4 min 10 s task is charged 250 seconds: 250 s × $5.40 / 3,600 s = $0.38.
| State | Compute charge |
|---|---|
| running | Spot price for each interval |
| stopping | Spot price until the node stops |
| stopped_price | None. Storage continues. |
| stopped_reclaimed | None. Storage continues. |
| resuming | None until running |
- Sheet
- 01 / 01
- Title
- Batch Inference: queue and billing
- Metering
- Per second
- Reviewed
- 2026-10-10
Questions about Batch Inference.
Which models can I run?
Anything that fits your container image and the GPU you choose. Batch Inference runs your image; it does not host a model catalog.
Where do inputs and outputs live?
In your object storage. You pass input and output URIs, and workers read and write there.
Do I pay my max price?
No. You pay the spot price for each interval. Your max price decides whether your nodes run.
What happens when the spot price goes above my max?
Those nodes get a 60-second notice and stop. Their disks are kept. Batch Inference and managed jobs resume from their last checkpoint when the price is back under your max.
What does Blocked mean?
Your max price is below the current spot price, so nothing has started. It starts when the price falls to your max, or when you raise it.
What happens if spot capacity disappears for hours?
Jobs wait in the queue at their priority. Low-priority jobs wait for spot capacity under your ceiling.
Can spot capacity stop even if I follow the market?
Yes. Following the market removes the price limit, but spot capacity can still be reclaimed when it is needed elsewhere. How spot works
Is spot capacity guaranteed?
No. Spot is best effort, and the SLA excludes it.