H200-SXMHopper
NVIDIA H200 SXM
Memory-bound training and long-context inference on the same HGX platform as H100.
HBM3E per GPU
141 GB
NVIDIA reference figure
- 1. SXM module · 8 per baseboard, in a 2 × 4 grid
- 2. Heatsink · one per module; this one is drawn lifted
- 3. H200 GPU · one per module, under the heatsink
- 4. NVSwitch chip · 4 per baseboard, grouped at one end
- 5. Baseboard · joins the 8 GPUs in one NVLink domain
Point at a part or a row. The other one lights.
H200 SXM specifications.
| Parameter | Value | Scope | Source |
|---|---|---|---|
| Memory | |||
| Memory | 141 GB HBM3E | Per GPU | [1] |
| Memory bandwidth | 4.8 TB/s | Per GPU | [1] |
| Memory on one board | 1,128 GB · 8 × 141 GB | Per HGX board | [1] |
| Compute, peak · dense / sparse | |||
| FP4 Tensor Core | Not supported on Hopper | Per GPU | |
| FP8 Tensor Core | 1,979 / 3,958 TFLOPS | Per GPU | [1] |
| BF16 Tensor Core | 989 / 1,979 TFLOPS | Per GPU | [1] |
| Interconnect | |||
| NVLink 4 | 900 GB/s | Per GPU | [1] |
| NVLink domain | 8 GPUs per HGX H200 board | Per HGX board | [2] |
| Host interface | PCIe Gen5 x16 · 128 GB/s | Per GPU | [1] |
| Power and form factor | |||
| Power | Up to 700 W, configurable | Per GPU | [1] |
| Form factor | SXM module on the HGX H200 baseboard | [1] | |
| Board | 8 GPUs per HGX H200 board · 4 NVSwitch chips (third generation) | Per HGX board | [2] |
Specifications are NVIDIA reference figures.
Best for
- Memory-bound LLM inference
- Training and fine-tuning
- HPC
Notes on the figures
- FP4: not supported on Hopper (FP4 Tensor Cores arrived with Blackwell).
- FP8 and BF16: NVIDIA marks the larger figures "with sparsity"; dense is half.
On-demand, spot or reserved.
Metered by the second, billed hourly. Reserved rates are quoted per term.
-
On-demand
$5.94 /GPU·hr
- Per GPU-second
- $0.001650
- Per 8 GPUs, per hour
- $47.52
- Per GPU-month, 730 h
- $4,336.20
Preview list price. It applies when your account opens.
-
Spot
From $0.87 /GPU·hr
Work that checkpoints and can wait. Follow the current spot price, or set a maximum hourly price. If the spot price rises above your max, the workload may stop.
"From" is the lowest spot price for that GPU. The spot price moves with supply and demand. The spot price always stays below the on-demand rate for the same GPU.
Spot price moves with supply and demand. Follow it, or set a max price per GPU-hour: you pay the spot price, and nodes stop if it rises above your max. A max price limits what you pay. It does not reserve capacity. To guarantee capacity, reserve it. How spot works
-
Reserved
Request a quote
Capacity you need on a date, guaranteed. 1 month, 6 months or 12+ months.
Ways to pay Prepaid credit, pay-as-you-go and Enterprise. Each one covers on-demand, spot and reserved capacity. How billing works
| Term | Commitment | Capacity | Rate |
|---|---|---|---|
| On-demandNone · Not held | None | Not held | $10.45 /GPU·hr |
| 1 monthTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 6 monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 12+ monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| Term | Commitment | Capacity | Rate |
|---|---|---|---|
| On-demandNone · Not held | None | Not held | $9.35 /GPU·hr |
| 1 monthTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 6 monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 12+ monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| Term | Commitment | Capacity | Rate |
|---|---|---|---|
| On-demandNone · Not held | None | Not held | $5.94 /GPU·hr |
| 1 monthTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 6 monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 12+ monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| Term | Commitment | Capacity | Rate |
|---|---|---|---|
| On-demandNone · Not held | None | Not held | $5.40 /GPU·hr |
| 1 monthTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 6 monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| 12+ monthsTake-or-pay · Held for the term, on one fabric | Take-or-pay | Held for the term, on one fabric | Request a quote |
| Term | Commitment | Capacity | Rate |
|---|---|---|---|
| On-demandNone · Not held | None | Not held | $2.16 /GPU·hr |
| 1 monthTake-or-pay · Held for the term | Take-or-pay | Held for the term | Request a quote |
| 6 monthsTake-or-pay · Held for the term | Take-or-pay | Held for the term | Request a quote |
| 12+ monthsTake-or-pay · Held for the term | Take-or-pay | Held for the term | Request a quote |
| Term | Commitment | Capacity | Rate |
|---|---|---|---|
| On-demandNone · Not held | None | Not held | $1.86 /GPU·hr |
| 1 monthTake-or-pay · Held for the term | Take-or-pay | Held for the term | Request a quote |
| 6 monthsTake-or-pay · Held for the term | Take-or-pay | Held for the term | Request a quote |
| 12+ monthsTake-or-pay · Held for the term | Take-or-pay | Held for the term | Request a quote |
List price for the base shape: 1 GPU, 8 vCPU, 32 GiB. Larger shapes are quoted.
Reserved rates are quoted per term. Reservations are tied to a GPU type, a region and one InfiniBand fabric. Take-or-pay under a signed order form. No early cancellation without penalty (Terms §3).
Runs as instances and as cluster nodes.
HGX H200: 8 GPUs, 4 third-generation NVSwitch chips, 1,128 GB of HBM3E per node.
H200 runs as a 1-GPU instance (16 vCPU · 200 GiB) or as an 8-GPU node (128 vCPU · 1,600 GiB). 8-GPU nodes join one InfiniBand fabric at 400 Gb/s per GPU, 3.2 Tb/s per node. Host network adapter: 200 Gb/s per machine. Connections: 8 GPUs to One InfiniBand fabric, 8 links: 8 × 400 Gb/s; More 8-GPU nodes to One InfiniBand fabric; 1 GPU to Host network adapter.
- vCPU
- 16
- RAM
- 200 GiB
1 GPU- rating
- 200 Gb/s
- scope
- per machine
- vCPU
- 128
- RAM
- 1,600 GiB
- CPU
- Intel Xeon Platinum
- 8468
8 GPUs · NVLink 4- per GPU
- 400 Gb/s
- per node
- 3.2 Tb/s
Every node of one cluster sits on the same fabric. A cluster is placed onone fabric and stays there.
- vCPU
- 16
- RAM
- 200 GiB
1 GPU- rating
- 200 Gb/s
- scope
- per machine
- vCPU
- 128
- RAM
- 1,600 GiB
- CPU
- Intel Xeon
- Platinum 8468
8 GPUs · NVLink 4- per GPU
- 400 Gb/s
- per node
- 3.2 Tb/s
- vCPU
- 16
- RAM
- 200 GiB
1 GPU- rating
- 200 Gb/s
- scope
- per machine
- vCPU
- 128
- RAM
- 1,600 GiB
- CPU
- Intel Xeon Platinum 8468
8 GPUs · NVLink 4- per GPU
- 400 Gb/s
- per node
- 3.2 Tb/s
H200 SXM listing.
CPU, RAM, local NVMe and the fabric vary by pool and are reported at placement. GPUs and InfiniBand adapters are passed through whole to one machine.
- Status
- Private Preview
- Runs as
- Single instances and GPU Clusters
- Capacity
- On-demand · Spot · Reserved
- Metering
- Metered by the second, billed hourly
| Parameter | Value | Scope |
|---|---|---|
| 1 GPU | 16 vCPU · 200 GiB | Per machine |
| 8 GPUs | 128 vCPU · 1,600 GiB | Per machine |
| Host CPU | Intel Xeon Platinum 8468 | Per 8-GPU node |
| InfiniBand | 400 Gb/s | Per GPU |
| InfiniBand | 3.2 Tb/s | Per 8-GPU node |
| Host network adapter | 200 Gb/s | Per machine |
NVIDIA H200 SXM on Fantasti: fantasti.ai/gpus/h200-sxm · Request capacity: fantasti.ai/contact
- Sheet
- 01 / 01
- Title
- H200 SXM datasheet
- Figures
- NVIDIA reference
- Reviewed
- 2026-10-10
Request H200 SXM capacity.
Tell us the count, the term and where it should run. We reply to a reviewed request, usually within one business day. A reserved term is quoted.