DevelopFlows Planned
From experiment to production, in one Python flow.
Flows runs your ML work as versioned Python workflows on Fantasti GPUs. Write each step as a function, give it the GPU it needs, and run the same file from a laptop, a workspace or a schedule. Every run keeps its code, data, checkpoints and results.
One run of FinetuneFlow, started with python finetune.py run: five steps in order. The train step runs on H200:8 on spot with a max price of 3.10 USD per GPU-hour, stops once when the spot price rises above the max and resumes from checkpoint 3. The evaluate step runs 12 eval cases on L40S:1, and the promote step compares the score 0.86 with the gate 0.82 and publishes the model. Connections: support_sft to start; start to train: data; train to evaluate: weights; evaluate to promote: scores; promote to end: passed.
# five steps, two of them on GPUs$ python finetune.py run- data
- s3://acme-data/tickets @ v41
done@fantasti(gpu="H200:8", capacity="spot",max_price=3.10)@checkpoint # resume hererunning max_price 3.10stopped price above maxresumed checkpoint 3HGX H200 · 8 GPUs
- 12 eval cases
done exact_match 0.86 - score 0.86gate 0.82
- endpoint
- support-sft
done promoted donerun_7c2m…
# five steps, two on GPUs$ python finetune.py run- data
- s3://acme-data/tickets @ v41
done@fantasti(gpu="H200:8", capacity="spot",max_price=3.10)@checkpoint # resume hererunning max_price 3.10stopped price above maxresumed checkpoint 3HGX H200 · 8 GPUs
- 12 eval cases
done exact_match 0.86 - score 0.86gate 0.82
- endpoint
- support-sft
done promoted donerun_7c2m…
# five steps, two of them on GPUs$ python finetune.py run- data
- s3://acme-data/tickets @ v41
done@fantasti(gpu="H200:8",capacity="spot",max_price=3.10)@checkpoint # resume hererunning max_price 3.10stopped price above maxresumed checkpoint 3HGX H200 · 8 GPUs
- 12 eval cases
done exact_match 0.86 - score 0.86gate 0.82
- endpoint
- support-sft
done promoted donerun_7c2m…
- This run
- Framework
- open-source Metaflow
- Specific to Fantasti
- one decorator per GPU step
- Capacity
- on-demand · spot with max_price
- Billing
- Compute rates · no platform fee
from metaflow import FlowSpec, Parameter, step, project, schedule, retry
from metaflow import card, checkpoint, model, current, fantasti
@project(name="support_sft")
@schedule(cron="0 6 * * 1") # Mondays 06:00 UTC, once deployed
class FinetuneFlow(FlowSpec):
@fantasti(gpu="H200:8", capacity="spot", max_price=3.10) # USD/GPU-hour
@checkpoint # after a spot stop, resume here
@model # version the weights with the run
@retry(times=4)
@step
def train(self):
out = sft(self.base, self.data, ckpt_dir=current.checkpoint.directory)
self.weights = current.model.save(out, label="support-sft")
self.next(self.evaluate)
# start, evaluate, promote and end are in the Evals section $ python finetune.py run # from a laptop or a workspace
$ python finetune.py --production fantasti create # deploy the schedule Output of run · illustrative
start local done
train H200:8 spot running max_price 3.10
train H200:8 spot stopped price above max
train H200:8 spot resumed checkpoint 3
evaluate L40S:1 done exact_match 0.86
promote local done promoted
end local done run_7c2m… // .mcp.json · project scope, OAuth sign-in
{
"mcpServers": {
"fantasti": { "type": "http", "url": "https://mcp.fantasti.ai/mcp" }
}
}
// Flows tools · planned
// flows_list · flows_trigger · runs_get
// runs_logs · runs_resume · artifacts_read
// asks before it acts: flows_trigger, runs_resume curl -X POST https://api.fantasti.ai/v1/flows/support_sft.prod.FinetuneFlow/runs \
-H "Authorization: Bearer $FANTASTI_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "parameters": { "gate": 0.85 } }'
# { "id": "run_7c2m…", "status": "placing", "trigger": "api" } # flows.yaml · defaults for every step in this project
project: support_sft
image: ghcr.io/acme/train:latest
capacity: { mode: spot, max_price: follow } # a step decorator overrides this
artifacts: s3://acme-flows/ # your bucket
mlflow: default # sets MLFLOW_TRACKING_URI in every step
secrets: [HF_TOKEN] It's standard Metaflow.
A flow is a Python class. Each step is a method, and anything you assign to self is stored and versioned when the step ends. Flows runs open-source Metaflow, so this file also runs on a laptop with no cluster at all.
One line is specific to Fantasti: the decorator that says which GPU a step needs and how to pay for it.
Metaflow is open-source software. Its use here does not imply endorsement by its maintainers.
1 line of this flow is specific to Fantasti.
-
@fantasti(gpu="H200:8", capacity="spot", max_price=3.10)Fantasti -
@checkpointMetaflow -
@modelMetaflow -
@retry(times=4)Metaflow -
@stepMetaflow -
def train(self):Python
- Write
- A Python class
- Steps are methods. State is whatever you assign to
self. - Run
- Laptop, workspace or schedule
- The same file in all three. Steps that ask for GPUs run on your cluster.
- Keep
- Every run, by id
- Code, parameters, data, checkpoints and results.
What Flows gives you.
Six things every flow has, from the run on your laptop to the run on a schedule.
Flows as code
A flow is a Python file in your repository. Review it, test it and branch it like the rest of your code.
- Run anywherethe same file runs on a laptop, in a workspace or on your cluster.
- Resumerestart a failed run from the step that failed.
- Brancheseach developer and each pull request gets its own copy of the flow.
The file finetune.py defines the class FinetuneFlow with one method per step. The command python finetune.py run runs it. Connections: finetune.py to cmd.
class FinetuneFlow(FlowSpec):@stepdef train(self):$ python finetune.py run
A GPU per step, at your price
Each step asks for what it needs: one L40S for evals, an 8-GPU H200 or B200 node for training.
- Spot with a ceilingfollow the market or set a maximum per GPU-hour.
- One fabrica step that spans nodes stays on one InfiniBand fabric.
- ReleasedGPUs are released when the step ends.
The train step runs on H200:8 on spot with a max price of 3.10 USD per GPU-hour. The evaluate step runs on L40S:1 on demand. Connections: train to evaluate.
- max 3.10
- on demand
Versioned by default
Every run records the code, parameters, container image and data it used, and every artifact it produced.
- Artifactsstored in your bucket, addressed by flow, run and step.
- Modelscheckpoints and weights are versioned with the run that made them.
- MLflowmetrics and model versions land in managed MLflow.
Run run_7c2m… records its code (git 4e1a9c2), its data (tickets @ v41) and its parameters (gate 0.82). Its artifacts are stored in the bucket s3://acme-flows/. Connections: run_7c2m… to s3://acme-flows/.
- code
- git 4e1a9c2
- data
- tickets @ v41
- params
- gate=0.82
Schedules and events
Deploy a flow and it runs without a laptop: on a schedule, when new data arrives, or when another flow finishes.
- Cronhourly, daily or any cron expression.
- Eventstrigger on a named event or an upstream flow.
- Isolatedproduction runs are separate from development runs of the same flow.
The deployed flow FinetuneFlow runs on its schedule, cron 0 6 * * 1, which is every Monday. A named event starts a run between two scheduled ones. Connections: Schedule to FinetuneFlow; Event to FinetuneFlow.
@schedule(cron="0 6 * * 1")
Evals as gates
An eval is a step. A gate is a comparison. A model is promoted when its scores pass.
- Parallelfan eval cases out across GPUs or sandboxes.
- Reportseach run renders its results as a report you can link to.
- Thresholdsset the scores a model must reach, per flow.
The evaluate step runs 12 eval cases. The promote step compares the score 0.86 with the gate 0.82, and the model passes. Connections: evaluate to promote.
scores["exact_match"] >= gatescore 0.86gate 0.82
Deploy and observe
A passing model goes where it is used, and the record says which run made it.
- Publishto a Serverless endpoint or a Batch Inference run.
- Watchlogs, GPU metrics and step timelines in the console.
- Costper run and per step, on your Fantasti statement.
The promote step publishes the model to the Serverless endpoint support-sft. The endpoint names the run that made it, run_7c2m…. Connections: promote to support-sft.
- product
- Serverless
- made by
- run_7c2m…
- score
- 0.86 · gate 0.82
Every run leaves a record.
Open any run and see what it used, where it ran and what it produced.
| Field | This run | The run before it |
|---|---|---|
| What it used | ||
| flow | FinetuneFlow · support_sft/prod | FinetuneFlow · support_sft/prod (same) |
| run | run_7c2m… Before run_6v1k… | run_6v1k… |
| trigger | schedule · Mon 06:00 UTC | schedule · Mon 06:00 UTC (same) |
| code | git 4e1a9c2 · image sha256:9f3b… | git 4e1a9c2 · image sha256:9f3b… (same) |
| data | s3://acme-data/tickets @ v41 Before s3://acme-data/tickets @ v40 | s3://acme-data/tickets @ v40 |
| Where it ran | ||
| train | H200:8 · spot · max_price 3.10 · resumed 1× Before H200:8 · spot · max_price 3.10 · resumed 0× | H200:8 · spot · max_price 3.10 · resumed 0× |
| evaluate | L40S:1 · on-demand | L40S:1 · on-demand (same) |
| What it produced | ||
| result | exact_match 0.86 · gate 0.82 · passed Before exact_match 0.79 · gate 0.82 · held | exact_match 0.79 · gate 0.82 · held |
| state | succeeded · published to endpoint support-sft Before succeeded · not published | succeeded · not published |
5 of 9 fields differ. The run before it (run_6v1k…) read data v40 and was held by the gate. Its values are marked and set in ink where the two runs differ.Its value is listed where the two runs differ. Trace both runs through the gate
- Sheet
- 01 / 03
- Title
- Run record
- Run
- run_7c2m…
- Stamp
- Illustrative
Spot for the long step. On-demand for the rest.
Each step carries its own capacity setting. Put the long training step on spot with a maximum price and keep the short steps on demand. While a spot step runs you pay the spot price, which is at or below your max.
If the spot price rises above your max, that step stops. Flows keeps the run open, and the step resumes from its last checkpoint when capacity at your price returns. Steps that already finished are not rerun.
The Fantasti Orchestrator places each step. It finds the GPU type you asked for and keeps a multi-node step on one InfiniBand fabric.
- A max price does not reserve capacity.
- Steps run on H100, H200, B200, B300, RTX PRO 6000 and L40S. GB200 NVL72 and GB300 NVL72 are reserved rack-scale capacity, not step shapes.
- Step
- train · H200:8
- Max price
- $3.10 /GPU·hr
- Stopped
- 1×
- Resumed from
- checkpoint 3
Illustrative series, not Fantasti market data. A step line shows a spot price over one run of FinetuneFlow, between $2.41 and $3.47 per GPU-hour. The train step runs on H200:8 on spot with a max price of $3.10. It writes checkpoints 1 to 3, stops once when the spot price rises above the max, resumes from checkpoint 3 when the price falls back, writes checkpoint 4 and finishes. The evaluate step then runs on L40S:1 on demand: the spot price rises above the max again while it runs, and it is not stopped.
From a laptop run to a schedule.
| Step | You | Fantasti |
|---|---|---|
| 01Write | Write the flow in your editor or a workspace. | Nothing yet. A flow runs locally with no cluster. |
| 02Run | python finetune.py run | Sends GPU steps to your cluster, streams logs back and stores every artifact. |
| 03Branch | Open a pull request. | Deploys the branch as its own copy of the flow, separate from production. |
| 04Deploy | Merge, then deploy the flow with one command.python finetune.py --production fantasti create | Runs it on its schedule or trigger. No laptop or workspace needs to stay on. |
| 05Promote | Set the scores a model must reach. | Runs the evals, records the result and publishes the model when it passes. |
- Sheet
- 02 / 03
- Title
- From a laptop run to a schedule
- Stamp
- Planned
# finetune.py, continued: the steps after train
class FinetuneFlow(FlowSpec):
gate = Parameter("gate", default=0.82) # minimum exact_match to promote
@fantasti(gpu="L40S:1")
@model(load="weights")
@card
@step
def evaluate(self):
self.scores = run_evals(current.model.loaded["weights"],
suite="evals/support-v3.jsonl")
self.next(self.promote)
@step
def promote(self):
self.passed = self.scores["exact_match"] >= self.gate
if self.passed:
publish(self.weights, endpoint="support-sft") # a Serverless endpoint
self.next(self.end) An eval is a step. A gate is a comparison.
Evals run inside the flow that trained the model, on the same data version, so the score and the weights share one record. Fan the cases out across GPUs, or run each case in a closed-network sandbox when the code under test is untrusted.
The promote step compares the scores with the threshold you set. A model that passes is registered in managed MLflow and published. A model that fails stays in the run record with the reason.
run_7c2m… read data v41 and scored 0.86. That is at or above the gate of 0.82, so the model is registered in managed MLflow and published to support-sft.
The evaluate step runs 12 eval cases on L40S:1 and hands its scores to the promote step, which compares exact_match with the gate of 0.82. run_7c2m… scored 0.86 and passed: its model is registered in managed MLflow and published to the Serverless endpoint support-sft. run_6v1k…, the scheduled run before it, scored 0.79 and was held: nothing is published, and the reason stays in its run record. Connections: evaluate to promote: scores; promote to support-sft; promote to run_6v1k….
- 12 eval cases
- suite
- evals/support-v3.jsonl
- data
- tickets @ v41
- data
- tickets @ v40
- score
- exact_match 0.86
- score
- exact_match 0.79
scores["exact_match"] >= gate- run_7c2m…
- 0.86 · passed
- run_6v1k…
- 0.79 · held
- score 0.86gate 0.82
- model
- from run_7c2m…
- registry
- managed MLflow
- state
- published
- score 0.79gate 0.82
- state
- not published
- reason
- 0.79 below gate 0.82
- 12 eval cases
- suite
- evals/support-v3.jsonl
- data
- tickets @ v41
- data
- tickets @ v40
- score
- exact_match 0.86
- score
- exact_match 0.79
scores["exact_match"] >= gate- run_7c2m…
- 0.86 · passed
- run_6v1k…
- 0.79 · held
- score 0.86gate 0.82
- model
- from run_7c2m…
- registry
- managed MLflow
- state
- published
- score 0.79gate 0.82
- state
- not published
- reason
- 0.79 below gate 0.82
- 12 eval cases
- suite
- evals/support-v3.jsonl
- data
- tickets @ v41
- data
- tickets @ v40
- score
- exact_match 0.86
- score
- exact_match 0.79
scores["exact_match"] >= gate- run_7c2m…
- 0.86 · passed
- run_6v1k…
- 0.79 · held
- to
- run_6v1k…
- score 0.86gate 0.82
- model
- from run_7c2m…
- registry
- managed MLflow
- state
- published
- score 0.79gate 0.82
- state
- not published
- reason
- 0.79 below gate 0.82
- The run you trace
- Every run
Flows, or something else?
| Question | Flows Planned | Workspaces Early access | Sandboxes Early access | Batch Inference | Serverless Early access |
|---|---|---|---|---|---|
| You bring | A Python flow | Your editor and code | An agent or test harness | A model and a dataset | A container |
| Fantasti runs | Each step, in order, on the GPU it names | A dev environment that scales to a cluster | Short-lived isolated environments | A sharded inference run | One job or one endpoint |
| Ends | When the last step does; then on its next trigger | When you stop it | At its time-to-live | When every shard is done | When the job exits or you stop the endpoint |
| Keeps | Every run, versioned | Your files | Nothing; copy results out | Outputs and checkpoints | Nothing on the container disk |
| Capacity | Per step: on-demand or spot | On-demand head; workers on-demand or spot | On-demand | Spot or on-demand | On-demand or spot |
| Use it when | The work repeats and must be reproducible | You are writing or debugging | Code is untrusted or disposable | The job is one model over one dataset | The job is one container |
- Flows and MLflow: Flows runs the work. MLflow keeps the record of metrics and model versions.
- Flows and Integrations: already run Ray, SkyPilot or NVIDIA OSMO? Keep it. Flows is for teams that want Fantasti to run the scheduler.
- Flows and the Orchestrator: Flows decides what runs and in what order. The Orchestrator decides where each step runs and what it costs.
You pay for the GPUs your steps use.
- Platform fee
- None
- No fee for Flows itself and no charge per seat.
- Steps
- Compute rates, per second
- Each step is metered from start to end at the rate of the GPU or CPU it uses.
- Statement
- Cost per run and per step
- On the same statement as the rest of your account.
| Step | Runs on | Capacity | Per GPU-hour | Per GPU-second | This flow sets |
|---|---|---|---|---|---|
| train | 8 × H200 SXM | Spot | From $0.87 | From $0.000242 | max_price 3.10 |
| evaluate | 1 × L40S | On-demand | $1.86 | $0.000517 | — |
Metered by the second, billed hourly. "From" is the lowest spot price for that GPU. The spot price moves with supply and demand. L40S List price for the base shape: 1 GPU, 8 vCPU, 32 GiB. Larger shapes are quoted. All prices
Questions and answers.
Do I have to learn a new framework?
Is my code tied to Fantasti?
What happens when a spot step is stopped?
Where do my artifacts live?
Do I still need MLflow?
Can a step use more than one node?
Can a coding agent run flows?
Is Fantasti affiliated with the Metaflow project?
Can I use Flows yet?
- Sheet
- 03 / 03
- Title
- Billing and questions
- Stamp
- Planned
Bring one flow. We'll put it on a schedule.
Tell us what you train and how often. We are choosing design partners.