Research

Research3 lines of work

How a GPU reaches a workload.

Today a cloud GPU is installed in one server and rented with it. Fantasti studies what changes when the GPU is attached over a network instead: how the calls travel, where the work is placed, and what happens when capacity is taken back.

Research in progress. Results are published as numbered notes, each with its method and hardware.

Research register. Schematic · not to scale. Reviewed 2026-10-09.

A program on a machine with no GPU, a GPU reached over the network, and the 3 lines of work on the path between them: R-01, Network-based GPU virtualization, exploring, 0 notes published; R-02, Delivery protocols for GPU workloads, exploring, 0 notes published; R-03, Placement and interruption prediction, exploring, 0 notes published. Connections: train.py to PCI Express slot; PCI Express slot to GPU; train.py to Network-based GPU virtualization: calls; Network-based GPU virtualization to Delivery protocols for GPU workloads; Delivery protocols for GPU workloads to Placement and interruption prediction; Placement and interruption prediction to GPU.

  • Programtrain.py
    device = torch.device("cuda")# this machine has no GPU
  • R-01Network-based GPU virtualization

    Which call patterns survive a network round trip,and which do not.

    Exploring · stage 1 of 4
  • R-02Delivery protocols for GPU workloads

    What a transport should guarantee when the payloadis GPU calls and tensors, not web requests.

    Exploring · stage 1 of 4
  • R-03Placement and interruption prediction

    How much warning is possible before capacity istaken back, and what a scheduler should do with anuncertain forecast.

    Exploring · stage 1 of 4
  • PCI Express slotToday
  • GPUOn the network

    Runs the callsof a program onanother machine.

  • Programtrain.py
    device = torch.device("cuda")# this machine has no GPU
  • R-01Network-based GPU virtualization
    Exploring · stage 1 of 4
  • R-02Delivery protocols for GPU workloads
    Exploring · stage 1 of 4
  • R-03Placement and interruption prediction
    Exploring · stage 1 of 4
  • GPUOn the network

    Runs the calls of a program on anothermachine.

  • Programtrain.py
    device = torch.device("cuda")# this machine has no GPU
  • R-01

    Network-based GPU virtualization

    Exploring · stage 1 of 4
  • R-02

    Delivery protocols for GPU workloads

    Exploring · stage 1 of 4
  • R-03

    Placement and interruption prediction

    Exploring · stage 1 of 4
  • GPUOn the network

    Runs the calls of aprogram on anothermachine.

  • The path we study
  • Today: in the same server
The lines, written out
Lines of work
3 · exploring
Reading list
10 works
Where we work
client · cluster · control plane
Every result states
hardware · baseline · spread
§01 Why

Compute, storage and networking were each separated from the server and pooled. The GPU is still rented attached to one machine. We are studying how far that can change, and what it costs.

  • Compute Pooled
  • Storage Pooled
  • Networking Pooled
  • GPU Attached to one server
§02 Attach point Source · §07 reading list

Three places to attach a GPU.

One is how clouds work today. One needs hardware at both ends of the link. One is software, and that is where we work.

Three places to attach a GPU. Schematic · not to scale.

The same stack drawn three times on five levels: program, GPU library, driver, PCI Express link and GPU. Local: everything is in one server, and another server keeps its own GPU, idle. Fabric: the PCI Express link crosses a switch fabric from a host adapter to a target adapter and the GPU in a separate chassis, and the server still sees a local device. API: a machine without a GPU holds the workload and a client library, and library calls cross the network to a server process on a GPU server, which holds the GPU library, the driver and the GPU. The API path is the one Fantasti works on. Connections: Workload to GPU library; GPU library to Driver; Driver to GPU: slot; Workload to GPU: no path; Driver to Host adapter; Host adapter to GPU; Host adapter to Target adapter; Target adapter to GPU; Workload to Client library; Server process to GPU library; Driver to GPU; Client library to Server process: calls.

ALocalHow clouds work today
Server
Another server
BFabricHardware at both ends
Server
GPU chassis
CAPISoftware · where we work
Machine · no GPU
GPU server
  • Workload
  • GPU library
  • Driver
  • GPU
  • GPUIdle
  • Workload
  • GPU library
  • Driver
  • Host adapter
  • GPUAs seen
  • Target adapter
  • GPU
  • Workload
  • Client library
  • Server process
  • GPU library
  • Driver
  • GPU
Attach pointPCI Express slotAttach pointSwitch fabricAttach pointLibrary call
  • Where we work
  • Library calls over the network
  • Call into the next layer
  • No path to another server
  • Seen as a local device
  • Idle

A. Local: how clouds work today

Local: where the GPU attaches. Schematic · not to scale.

One server holds the workload, the GPU library, the driver and the GPU. The attach point is a PCI Express slot inside the server. Another server beside it keeps its own GPU, idle. Connections: Workload to GPU library; GPU library to Driver; Driver to GPU: slot; Workload to GPU: no path.

ALocalHow clouds work today
Server
Another server
  • Workload
  • GPU library
  • Driver
  • GPU
  • GPUIdle
Attach pointPCI Express slot
  • Workload
    to
    GPU
  • GPU library
  • Driver
  • GPU
  • GPUIdle
A · LocalServerAnother server
  • Call into the next layer
  • No path to another server
  • Idle

Where the GPU attachesThe GPU is a PCIe device inside the server. A virtual machine or container gets the whole device or a fixed slice of it.

What is knownThe baseline. Every framework and tool works. An idle GPU stays tied to its server.

B. Fabric: hardware at both ends

Fabric: where the GPU attaches. Schematic · not to scale.

The server holds the workload, the GPU library, the driver and a host adapter, and still sees a local GPU. The PCI Express link crosses a switch fabric to a target adapter and the GPU in a separate chassis. The attach point is the switch fabric. Connections: Workload to GPU library; GPU library to Driver; Driver to Host adapter; Host adapter to GPU; Host adapter to Target adapter; Target adapter to GPU.

BFabricHardware at both ends
Server
GPU chassis
  • Workload
  • GPU library
  • Driver
  • Host adapter
  • GPUAs seen
  • Target adapter
  • GPU
Attach pointSwitch fabric
  • Workload
  • GPU library
  • Driver
  • Host adapter
    to
    Target adapter
  • GPUAs seen
  • Target adapter
  • GPU
B · FabricServerGPU chassis
  • Call into the next layer
  • Seen as a local device

Where the GPU attachesThe PCIe link is carried across a switch fabric or wrapped in network packets. The server still sees a local device.

What is knownTransparent to software. It needs hardware at both ends, which the data-center operator controls.

C. API: software, where we work

API: where the GPU attaches. Schematic · not to scale.

A machine without a GPU holds the workload and a client library. Library calls cross the network to a server process on a GPU server, which holds the GPU library, the driver and the GPU. The attach point is the library call. This is the path Fantasti works on. Connections: Workload to Client library; Server process to GPU library; GPU library to Driver; Driver to GPU; Client library to Server process: calls.

CAPISoftware · where we work
Machine · no GPU
GPU server
  • Workload
  • Client library
  • Server process
  • GPU library
  • Driver
  • GPU
Attach pointLibrary call
  • Workload
  • Client library
  • Server process
  • GPU library
  • Driver
  • GPU
C · APIMachine · no GPUGPU server
  • Where we work
  • Library calls over the network
  • Call into the next layer

Where the GPU attachesGPU library calls are intercepted next to the workload and executed on a GPU somewhere else on the network.

What is knownNo special hardware. Compatibility and latency depend on how many calls cross the network and which features the workload uses.

Our research sits in the layers we run: the client library, the tenant cluster and the control plane.

§03 Lines of work Reviewed 2026-10-09
  1. R-01 Exploring , stage 1 of 4

    Network-based GPU virtualization

    Run an unmodified GPU program on a machine that has no GPU, with its calls executed on a GPU reached over the data-center network.

    Question
    Which call patterns survive a network round trip, and which do not.
    Boundary
    Single-GPU workloads inside one data center. Multi-GPU training with collective communication is out of scope for now.
    Expected to be hard
    Unified memory, profilers and other tools that assume a local device.
    We will publish
    A compatibility matrix by framework and feature, with the failures listed.
    torch.device("cuda")
    The aim: the program does not change. Where the call runs does.
  2. R-02 Exploring , stage 1 of 4

    Delivery protocols for GPU workloads

    Design the protocol between a workload and a remote GPU: how a session starts, how calls are batched, how memory moves, and how a GPU is detached and attached again.

    Question
    What a transport should guarantee when the payload is GPU calls and tensors, not web requests.
    Boundary
    Software protocols on standard Ethernet and IP, inside one data center.
    Expected to be hard
    Many small calls between large transfers, ordering on a byte stream, and state that must survive a detach.
    We will publish
    The message format and state machine as a written specification, before any performance number.
  3. R-03 Exploring , stage 1 of 4

    Placement and interruption prediction

    Decide where a workload should run, and estimate how long the spot capacity under it is likely to last, from price and availability history.

    Question
    How much warning is possible before capacity is taken back, and what a scheduler should do with an uncertain forecast.
    Boundary
    A forecast is an estimate. It can inform placement and checkpoint timing. It is not a guarantee and is not sold as one.
    Expected to be hard
    Short histories, markets that change behavior, and the cost of being wrong in either direction.
    We will publish
    Forecast error against held-out history, compared with a policy that uses no forecast, including the periods where the model was wrong.
§04 One call Illustrative
One call 5 lanes · 3 callouts
One GPU call, sent to a remote GPU. Illustrative.

One call across five lanes: Workload, Client library, Network, GPU server, GPU. The first two are on a machine without a GPU, the last two on the machine with the GPU. The workload makes a GPU call; the client library packs the arguments and sends a request across the network; the GPU server unpacks it and makes the driver call; the GPU executes; the GPU server packs the result and sends the reply; the client library unpacks the result and returns it. The request and the reply are the round trip, the packed arguments and results are the copy, and the GPU server keeps session state between calls. Order only: nothing is drawn to a time scale. Connections: train.py to GPU: what the program sees; train.py to Client library: GPU call; Client library to GPU server: request; GPU server to GPU: driver call; GPU to GPU server: result; GPU server to Client library: reply; Client library to train.py: return.

Machine without a GPU
Machine with the GPU
NetworkRound trip
  • Workloadtrain.py
    device = torch.device("cuda")
    01 GPU call
    a function call
    Result
    returned
  • Client library
    02 Pack arguments
    copy
    Runs
    beside the workload
    06 Unpack result
  • GPU serverA second process
    03 Unpack request
    Session state
    state
    05 Pack result
    copy
  • GPU
    04 Execute
  • Workloadtrain.py
    device = torch.device("cuda")
    01 GPU call
    a function call
    Result
    returned
  • Client library
    02 Pack arguments
    copy
    Runs
    beside the workload
    06 Unpack result
  • GPU server
    03 Unpack request
    Session state
    state
    05 Pack result
    copy
  • GPU
    04 Execute
Machine without a GPUMachine with the GPU
  • Where the call is intercepted
  • Crosses the network
  • A call inside one process
  • The call the program sees
§05 Method 6 rules
How results will be reported
No.Rule
01 Every number comes with its hardware, driver and framework versions, workload and baseline.
02 The baseline is the same workload on a locally attached GPU of the same model.
03 We report the spread across runs, not a single run.
04 Failures and unsupported features are listed in the same note as the results.
05 Notes are numbered, dated and versioned. Corrections are recorded in the note.
06 A research result does not become a product claim until a product ships with it.
§06 Notes Reviewed 2026-10-09

Notes.

Short technical notes, numbered and dated: problem statements, designs, methods and results, including the results that did not go our way.

Notes published

0

Kinds of note

  • Problem
  • Design
  • Method
  • Result
  • Negative result
  • Reading

A result note carries its hardware, its baseline and its method. Without them the site does not build.

TAB 02Notes register Source · Fantasti Reviewed 2026-10-11
Notes register. No notes yet.
No.DateTitleLineKindRev.

No notes yet. The register opens when there is something to report: a problem statement, a design or a measurement with its method. Until then this table stays empty on purpose.

§07 Prior work 10 works
TAB 03Reading list Source · the works cited
Reading list
Work Origin One line
rCUDA Universitat Politècnica de València, 2010 onward CUDA calls forwarded to GPUs on other cluster nodes, over TCP or InfiniBand.
GVirtuS Euro-Par 2010 GPU access from virtual machines through a forwarded API.
AvA ASPLOS 2020 Remoting stacks generated from an API description, with the hypervisor in the path.
DGSF 2022 Remote GPUs for serverless functions, with call batching and migration between GPUs.
Cricket RWTH Aachen, 2022 CUDA remoting with checkpoint and restart.
DxPU ACM TACO 2023 PCIe packets carried over the data-center network between hosts and GPU boxes.
Homa SIGCOMM 2018 A message-based, receiver-driven transport for data-center RPC.
SRD IEEE Micro 2020 Reliable datagrams delivered out of order across many paths.
Ultra Ethernet 1.0 Ultra Ethernet Consortium, 2025 An open Ethernet transport stack for AI and HPC.
Can't Be Late NSDI 2024 Policies for mixing spot and on-demand capacity under a deadline, with public availability traces.
Sheet
01 / 01
Title
Method, notes and reading list
Notes
0
Reviewed
2026-10-09
§08 Contact Research

If you work on GPU virtualization, data-center transports or scheduling, we would like to compare notes.