RunBatch Inference Private Preview

Stopped,
not lost.

Run async inference, evaluation and data-processing jobs on spot capacity. Workers checkpoint as they go, retry with backoff, and resume on new capacity when a node is reclaimed.

fx.batch.create(input="s3://…")

Opens with the first cohort. No account is open yet.

One run through a spot stop. Illustrative.

The run is queued, placed and runs on spot capacity with a max price. The spot price rises above the max: the nodes get a stop notice and stop, and their disks are kept. When the price is back at or below the max the run resumes, restores checkpoint 0412, runs again and completes. Inputs, checkpoints and results stay in your own bucket, s3://acme-evals/: workers read prompts/*.jsonl, write a checkpoint to ckpt/ every 5m, restore checkpoint 0412 from it after the stop, and write results to results/. Connections: Your request to Queue; Queue to bat_2kq8x1: placed; prompts/ to bat_2kq8x1: read; bat_2kq8x1 to ckpt/: write; ckpt/ to bat_2kq8x1: restore; bat_2kq8x1 to results/: results.

s3://acme-evals/Your bucket
  • PythonYour request
    fx.batch.create( gpu="H100:8", capacity=fantasti.Spot(max_price=2.40), checkpoint={"every": "5m"}, priority="low")
  • Queueyour quota
    order
    06 of 06
    waits for
    spot ≤ 2.40
    queued priority="low"
  • Runbat_2kq8x1H100:8 · spot ≤ 2.40
    placed us-central · fabric f · spotrunning ckpt 0001…0411stopped price above max · stop notice · disks keptresuming price back at or below maxrestored checkpoint 0412running ckpt 0413…0626
    HGX H100 8-GPU baseboard, line drawingHGX H100 · 8 GPUs
  • prompts/input
    *.jsonl
  • ckpt/every 5m
    restored
    0412
  • results/output
    done
s3://acme-evals/Your bucket
  • PythonYour request
    fx.batch.create( gpu="H100:8", capacity=fantasti.Spot(max_price=2.40), checkpoint={"every": "5m"}, priority="low")
  • Queue (done)
    order
    06 of 06
    max price
    2.40
    queued
  • Runbat_2kq8x1H100:8 · spot
    placed us-central · fabric f · spotrunning ckpt 0001…0411stopped price above maxresuming price back at or below maxrestored checkpoint 0412running ckpt 0413…0626
    HGX H100 8-GPU baseboard, line drawingHGX H100 · 8 GPUs
  • prompts/input
    *.jsonl
  • ckpt/every 5m
    restored
    0412
  • results/output
    done
  • PythonYour request
    fx.batch.create( gpu="H100:8", capacity=fantasti.Spot( max_price=2.40), checkpoint={"every": "5m"}, priority="low")
  • Queueyour quota
    order
    06 of 06
    waits for
    spot ≤ 2.40
    queued priority="low"
  • Runbat_2kq8x1H100:8 · spot
    placed us-central · fabric f · spotrunning ckpt 0001…0411stopped price above maxresuming price back at or below maxrestored checkpoint 0412running ckpt 0413…0626
    HGX H100 8-GPU baseboard, line drawingHGX H100 · 8 GPUs
    to
    results/
  • ckpt/
    written
    every 5m
    restored
    0412
  • results/output
    done
Your buckets3://acme-evals/
  • Restore from the checkpoint
  • This run
  • Reads and writes
bat_2kq8x1done

§01 Mechanics 4 parts

Mechanics

Checkpoint
checkpoint={"every": "5m"}
Checkpointing
Workers write progress at an interval you set and resume from the last checkpoint after a stop.
Retry
retries={"max": 3, …}
Backoff retries
Failed or stopped work retries with exponential backoff, up to the limit you set.
Worker
stopped → resuming → running
Workers that survive a stop
A worker that is reclaimed loses at most the work since its last checkpoint.
Queue
priority="low"
Queue priorities
Urgent jobs run first. Low-priority jobs wait for capacity under your price ceiling.
§02 API Preview API

Submit a job.
Collect results
in your bucket.

Inputs and outputs live in your object storage. You pass URIs; workers read and write there.

  • TrackedRuns log to managed MLflow. Early access
  • From a flowA flow step can start a batch run and wait for it. Planned
  • From your agentBatch tools on the MCP server. Planned
Preview API · subject to change
import fantasti

fx = fantasti.Client()  # reads FANTASTI_API_KEY

run = fx.batch.create(
    image="ghcr.io/acme/llm-eval:2026.10",   # your container
    input="s3://acme-evals/prompts/*.jsonl",
    output="s3://acme-evals/results/",
    gpu="H100:8",
    capacity=fantasti.Spot(max_price=2.40),  # USD per GPU-hour
    checkpoint={"every": "5m"},              # the resume point
    retries={"max": 3, "backoff": "exponential"},
    priority="low",                          # high, normal, low
)
print(run.id, run.status)  # bat_2kq8x1 queued
$ fantasti batch create \
    --image ghcr.io/acme/llm-eval:2026.10 \
    --input "s3://acme-evals/prompts/*.jsonl" \
    --output s3://acme-evals/results/ \
    --gpu H100:8 --spot --max-price 2.40 \
    --checkpoint 5m --retries 3 --priority low
$ fantasti batch get bat_2kq8x1 --watch

Output · illustrative

queued     priority="low"
placed     us-central · fabric f · spot
running    ckpt 0001…0411
stopped    price above max · stop notice · disks kept
resuming   price back at or below max
restored   checkpoint 0412 · s3://acme-evals/ckpt/
running    ckpt 0413…0626
done       s3://acme-evals/results/
// .mcp.json · project scope, OAuth sign-in
{
  "mcpServers": {
    "fantasti": {
      "type": "http",
      "url": "https://mcp.fantasti.ai/mcp"
    }
  }
}

Batch tools · planned

batch_create
batch_get
batch_cancel    asks for confirmation
curl -X POST https://api.fantasti.ai/v1/batch \
  -H "Authorization: Bearer $FANTASTI_API_KEY" \
  -H "Idempotency-Key: evals-1009" \
  -H "Content-Type: application/json" \
  -d '{
        "image": "ghcr.io/acme/llm-eval:2026.10",
        "input": "s3://acme-evals/prompts/*.jsonl",
        "output": "s3://acme-evals/results/",
        "gpu": "H100:8",
        "capacity": { "mode": "spot", "max_price": "2.40",
                      "currency": "USD", "unit": "gpu_hour" },
        "checkpoint": { "every": "5m" },
        "retries": { "max": 3, "backoff": "exponential" },
        "priority": "low"
      }'
# { "id": "bat_2kq8x1", "status": "queued" }
§03 Spot Illustrative series

Set a ceiling.
Run while the market is under it.

PLT 01Spot price and a max price over 72 hours Illustrative series · not Fantasti market data

Drag the max price line, click the plot, or use the arrow keys.

Run
H100:8
Max price
$2.40 /GPU·hr
Ran
51 h 15 m of 72 h
Stopped
4×
Resumed
4×

Illustrative series, not Fantasti market data. A step line shows a spot price for one GPU type over 72 hours in 15-minute steps, between $1.91 and $2.72 per GPU-hour. With the max price at $2.40, the nodes run 51 h 15 m of 72 hours, stop 4 times when the price rises above the max, and resume 4 times when it falls back. You pay the spot price for each interval, not the max.

EQ What a stop costs checkpoint every 5m

1 checkpoint interval is the most a reclaimed worker loses.

Work at risk when nodes stop
EQ 01 Source · Batch Inference API preview

In the illustrative run at the top of this page the nodes stop once. The run restores checkpoint 0412 and carries on from there. Nothing before it runs again.

DWG 02One stop, one interval repeated Illustrative

Illustrative. A worker writes a checkpoint at every interval (every 5m in the sample on this page). Its nodes stop partway through the interval after checkpoint 0412. After the run resumes, the worker restores that checkpoint and repeats only the work since it, which is at most one interval. Everything before the last checkpoint is kept.

Kept
Everything up to checkpoint 0412
Repeated
At most one interval (5m in this sample)
Compute charged while stopped
None. Storage continues.
§04 Priorities 3 priorities

Queue priorities.

TAB 01What each priority schedules Source · Fantasti Reviewed 2026-10-10
What each priority schedules
PrioritySchedulingUse for
highScheduled first, on any capacity your policy allowsRelease-blocking evals
normalScheduled after high-priority work, on spot capacity by defaultDaily inference and scoring
lowWaits for spot capacity under your price ceilingBackfills, re-scoring, data processing
TAB 02One queue, in scheduling order Illustrative

Illustrative. 6 runs from one organisation, in scheduling order, against a quota of 48 H100 GPUs. release-evals (high, 16 GPUs), hotfix-evals (high, 8 GPUs), daily-scoring (normal, 24 GPUs) fit inside the quota and are running. embed-docs (normal, 16 GPUs) waits for quota; backfill-q3 (low, 16 GPUs) waits for spot under max; rescore-evals (low, 8 GPUs) waits for spot under max.

One queue, in scheduling order
OrderRunPriorityCapacityGPUsState
01release-evals (bat_9c1r4m)highon-demand16running
02hotfix-evals (bat_7m3d0k)highspot ≤ 2.408running
03daily-scoring (bat_5t0w2p)normalspot ≤ 2.4024running
04embed-docs (bat_3f6q8a)normalspot ≤ 2.4016queued, waits for quota
05backfill-q3 (bat_1x8n5c)lowspot ≤ 2.4016queued, waits for spot under max
06rescore-evals (bat_2kq8x1)lowspot ≤ 2.408queued, waits for spot under max

One organisation, one quota of 48 H100 GPUs. Each bar is the GPUs a run asks for, placed after the runs scheduled ahead of it. Bars inside the quota are running; bars past it are queued.

§05 Billing Fantasti Terms §3

Pay for worker seconds.

Workers are metered by the second at the rate of the capacity they run on: the on-demand list price, or the spot price in effect for each interval. Time a worker spends stopped is not running and is not billed.

  • MeterMetered by the second, billed hourly.
  • On demand$5.40 per GPU-hour for H100 SXM5, the list price.
  • SpotFrom $0.87 per GPU-hour for H100 SXM5. You pay the spot price of each interval, not your max.

A stopped worker keeps its disks, so its storage continues.

EQ 02One 4 min 10 s task, billed three ways Source · Fantasti Terms §3
One 4 min 10 s task, billed three ways, at the H100 SXM5 preview list price of $5.40 per GPU-hour.
GranularitySeconds chargedFormulaCharge
Per hour3,600 s3,600 s × $5.40 / 3,600 s = $5.40$5.40
Per minute300 s300 s × $5.40 / 3,600 s = $0.45$0.45
Per second250 s250 s × $5.40 / 3,600 s = $0.38$0.38

Fantasti meters by the second and bills hourly (Terms §3). A 4 min 10 s task is charged 250 seconds: 250 s × $5.40 / 3,600 s = $0.38.

TAB 03Compute charge of a spot worker, by state Source · Fantasti Reviewed 2026-10-10
Compute charge of a spot worker, by state
StateCompute charge
runningSpot price for each interval
stoppingSpot price until the node stops
stopped_priceNone. Storage continues.
stopped_reclaimedNone. Storage continues.
resumingNone until running
Sheet
01 / 01
Title
Batch Inference: queue and billing
Metering
Per second
Reviewed
2026-10-10
§06 Questions 8 answers

Questions about Batch Inference.

Which models can I run?

Anything that fits your container image and the GPU you choose. Batch Inference runs your image; it does not host a model catalog.

Where do inputs and outputs live?

In your object storage. You pass input and output URIs, and workers read and write there.

Do I pay my max price?

No. You pay the spot price for each interval. Your max price decides whether your nodes run.

What happens when the spot price goes above my max?

Those nodes get a 60-second notice and stop. Their disks are kept. Batch Inference and managed jobs resume from their last checkpoint when the price is back under your max.

What does Blocked mean?

Your max price is below the current spot price, so nothing has started. It starts when the price falls to your max, or when you raise it.

What happens if spot capacity disappears for hours?

Jobs wait in the queue at their priority. Low-priority jobs wait for spot capacity under your ceiling.

Can spot capacity stop even if I follow the market?

Yes. Following the market removes the price limit, but spot capacity can still be reclaimed when it is needed elsewhere. How spot works

Is spot capacity guaranteed?

No. Spot is best effort, and the SLA excludes it.