Research3 lines of work
How a GPU reaches a workload.
Today a cloud GPU is installed in one server and rented with it. Fantasti studies what changes when the GPU is attached over a network instead: how the calls travel, where the work is placed, and what happens when capacity is taken back.
A program on a machine with no GPU, a GPU reached over the network, and the 3 lines of work on the path between them: R-01, Network-based GPU virtualization, exploring, 0 notes published; R-02, Delivery protocols for GPU workloads, exploring, 0 notes published; R-03, Placement and interruption prediction, exploring, 0 notes published. Connections: train.py to PCI Express slot; PCI Express slot to GPU; train.py to Network-based GPU virtualization: calls; Network-based GPU virtualization to Delivery protocols for GPU workloads; Delivery protocols for GPU workloads to Placement and interruption prediction; Placement and interruption prediction to GPU.
device = torch.device("cuda")# this machine has no GPUWhich call patterns survive a network round trip,and which do not.
Exploring · stage 1 of 4What a transport should guarantee when the payloadis GPU calls and tensors, not web requests.
Exploring · stage 1 of 4How much warning is possible before capacity istaken back, and what a scheduler should do with anuncertain forecast.
Exploring · stage 1 of 4Runs the callsof a program onanother machine.
device = torch.device("cuda")# this machine has no GPU- Exploring · stage 1 of 4
- Exploring · stage 1 of 4
- Exploring · stage 1 of 4
Runs the calls of a program on anothermachine.
device = torch.device("cuda")# this machine has no GPUNetwork-based GPU virtualization
Exploring · stage 1 of 4Delivery protocols for GPU workloads
Exploring · stage 1 of 4Placement and interruption prediction
Exploring · stage 1 of 4Runs the calls of aprogram on anothermachine.
- The path we study
- Today: in the same server
- Lines of work
- 3 · exploring
- Reading list
- 10 works
- Where we work
- client · cluster · control plane
- Every result states
- hardware · baseline · spread
Compute, storage and networking were each separated from the server and pooled. The GPU is still rented attached to one machine. We are studying how far that can change, and what it costs.
- Compute Pooled
- Storage Pooled
- Networking Pooled
- GPU Attached to one server
Three places to attach a GPU.
One is how clouds work today. One needs hardware at both ends of the link. One is software, and that is where we work.
The same stack drawn three times on five levels: program, GPU library, driver, PCI Express link and GPU. Local: everything is in one server, and another server keeps its own GPU, idle. Fabric: the PCI Express link crosses a switch fabric from a host adapter to a target adapter and the GPU in a separate chassis, and the server still sees a local device. API: a machine without a GPU holds the workload and a client library, and library calls cross the network to a server process on a GPU server, which holds the GPU library, the driver and the GPU. The API path is the one Fantasti works on. Connections: Workload to GPU library; GPU library to Driver; Driver to GPU: slot; Workload to GPU: no path; Driver to Host adapter; Host adapter to GPU; Host adapter to Target adapter; Target adapter to GPU; Workload to Client library; Server process to GPU library; Driver to GPU; Client library to Server process: calls.
- Where we work
- Library calls over the network
- Call into the next layer
- No path to another server
- Seen as a local device
- Idle
A. Local: how clouds work today
One server holds the workload, the GPU library, the driver and the GPU. The attach point is a PCI Express slot inside the server. Another server beside it keeps its own GPU, idle. Connections: Workload to GPU library; GPU library to Driver; Driver to GPU: slot; Workload to GPU: no path.
- to
- GPU
- Call into the next layer
- No path to another server
- Idle
Where the GPU attachesThe GPU is a PCIe device inside the server. A virtual machine or container gets the whole device or a fixed slice of it.
What is knownThe baseline. Every framework and tool works. An idle GPU stays tied to its server.
B. Fabric: hardware at both ends
The server holds the workload, the GPU library, the driver and a host adapter, and still sees a local GPU. The PCI Express link crosses a switch fabric to a target adapter and the GPU in a separate chassis. The attach point is the switch fabric. Connections: Workload to GPU library; GPU library to Driver; Driver to Host adapter; Host adapter to GPU; Host adapter to Target adapter; Target adapter to GPU.
- to
- Target adapter
- Call into the next layer
- Seen as a local device
Where the GPU attachesThe PCIe link is carried across a switch fabric or wrapped in network packets. The server still sees a local device.
What is knownTransparent to software. It needs hardware at both ends, which the data-center operator controls.
C. API: software, where we work
A machine without a GPU holds the workload and a client library. Library calls cross the network to a server process on a GPU server, which holds the GPU library, the driver and the GPU. The attach point is the library call. This is the path Fantasti works on. Connections: Workload to Client library; Server process to GPU library; GPU library to Driver; Driver to GPU; Client library to Server process: calls.
- Where we work
- Library calls over the network
- Call into the next layer
Where the GPU attachesGPU library calls are intercepted next to the workload and executed on a GPU somewhere else on the network.
What is knownNo special hardware. Compatibility and latency depend on how many calls cross the network and which features the workload uses.
Our research sits in the layers we run: the client library, the tenant cluster and the control plane.
-
R-01 Exploring , stage 1 of 4
Network-based GPU virtualization
Run an unmodified GPU program on a machine that has no GPU, with its calls executed on a GPU reached over the data-center network.
- Question
- Which call patterns survive a network round trip, and which do not.
- Boundary
- Single-GPU workloads inside one data center. Multi-GPU training with collective communication is out of scope for now.
- Expected to be hard
- Unified memory, profilers and other tools that assume a local device.
- We will publish
- A compatibility matrix by framework and feature, with the failures listed.
torch.device("cuda")The aim: the program does not change. Where the call runs does. -
R-02 Exploring , stage 1 of 4
Delivery protocols for GPU workloads
Design the protocol between a workload and a remote GPU: how a session starts, how calls are batched, how memory moves, and how a GPU is detached and attached again.
- Question
- What a transport should guarantee when the payload is GPU calls and tensors, not web requests.
- Boundary
- Software protocols on standard Ethernet and IP, inside one data center.
- Expected to be hard
- Many small calls between large transfers, ordering on a byte stream, and state that must survive a detach.
- We will publish
- The message format and state machine as a written specification, before any performance number.
-
R-03 Exploring , stage 1 of 4
Placement and interruption prediction
Decide where a workload should run, and estimate how long the spot capacity under it is likely to last, from price and availability history.
- Question
- How much warning is possible before capacity is taken back, and what a scheduler should do with an uncertain forecast.
- Boundary
- A forecast is an estimate. It can inform placement and checkpoint timing. It is not a guarantee and is not sold as one.
- Expected to be hard
- Short histories, markets that change behavior, and the cost of being wrong in either direction.
- We will publish
- Forecast error against held-out history, compared with a policy that uses no forecast, including the periods where the model was wrong.
One call across five lanes: Workload, Client library, Network, GPU server, GPU. The first two are on a machine without a GPU, the last two on the machine with the GPU. The workload makes a GPU call; the client library packs the arguments and sends a request across the network; the GPU server unpacks it and makes the driver call; the GPU executes; the GPU server packs the result and sends the reply; the client library unpacks the result and returns it. The request and the reply are the round trip, the packed arguments and results are the copy, and the GPU server keeps session state between calls. Order only: nothing is drawn to a time scale. Connections: train.py to GPU: what the program sees; train.py to Client library: GPU call; Client library to GPU server: request; GPU server to GPU: driver call; GPU to GPU server: result; GPU server to Client library: reply; Client library to train.py: return.
device = torch.device("cuda")- 01 GPU call
- a function call
- Result
- returned
- 02 Pack arguments
- copy
- Runs
- beside the workload
- 06 Unpack result
- 03 Unpack request
- Session state
- state
- 05 Pack result
- copy
- 04 Execute
device = torch.device("cuda")- 01 GPU call
- a function call
- Result
- returned
- 02 Pack arguments
- copy
- Runs
- beside the workload
- 06 Unpack result
- 03 Unpack request
- Session state
- state
- 05 Pack result
- copy
- 04 Execute
- Where the call is intercepted
- Crosses the network
- A call inside one process
- The call the program sees
| No. | Rule |
|---|---|
| 01 | Every number comes with its hardware, driver and framework versions, workload and baseline. |
| 02 | The baseline is the same workload on a locally attached GPU of the same model. |
| 03 | We report the spread across runs, not a single run. |
| 04 | Failures and unsupported features are listed in the same note as the results. |
| 05 | Notes are numbered, dated and versioned. Corrections are recorded in the note. |
| 06 | A research result does not become a product claim until a product ships with it. |
Notes.
Short technical notes, numbered and dated: problem statements, designs, methods and results, including the results that did not go our way.
Notes published
0
Kinds of note
- Problem
- Design
- Method
- Result
- Negative result
- Reading
A result note carries its hardware, its baseline and its method. Without them the site does not build.
| No. | Date | Title | Line | Kind | Rev. |
|---|---|---|---|---|---|
No notes yet. The register opens when there is something to report: a problem statement, a design or a measurement with its method. Until then this table stays empty on purpose.
| Work | Origin | One line |
|---|---|---|
| rCUDA | Universitat Politècnica de València, 2010 onward | CUDA calls forwarded to GPUs on other cluster nodes, over TCP or InfiniBand. |
| GVirtuS | Euro-Par 2010 | GPU access from virtual machines through a forwarded API. |
| AvA | ASPLOS 2020 | Remoting stacks generated from an API description, with the hypervisor in the path. |
| DGSF | 2022 | Remote GPUs for serverless functions, with call batching and migration between GPUs. |
| Cricket | RWTH Aachen, 2022 | CUDA remoting with checkpoint and restart. |
| DxPU | ACM TACO 2023 | PCIe packets carried over the data-center network between hosts and GPU boxes. |
| Homa | SIGCOMM 2018 | A message-based, receiver-driven transport for data-center RPC. |
| SRD | IEEE Micro 2020 | Reliable datagrams delivered out of order across many paths. |
| Ultra Ethernet 1.0 | Ultra Ethernet Consortium, 2025 | An open Ethernet transport stack for AI and HPC. |
| Can't Be Late | NSDI 2024 | Policies for mixing spot and on-demand capacity under a deadline, with public availability traces. |
- Sheet
- 01 / 01
- Title
- Method, notes and reading list
- Notes
- 0
- Reviewed
- 2026-10-09
If you work on GPU virtualization, data-center transports or scheduling, we would like to compare notes.
Next
- Orchestrator Policy, placement, pricing and recovery for every workload.
- Platform Every product on one account.
- Agents Designed to be operated by AI agents, under written rules and hard limits.
Who reads your request
The owner reads every request for the first cohort.