GB200-NVL72Blackwell (Grace Blackwell)
NVIDIA GB200 NVL72
72 Blackwell GPUs and 36 Grace CPUs in one liquid-cooled rack and one NVLink domain, on reserved terms.
In one NVLink domain
72 GPUs
NVIDIA reference figure
- 1. Power shelf · 8 per rack, six 5.5 kW supplies each
- 2. Management switch · 2 per rack
- 3. Compute tray · 18 × 1RU, 2 Grace CPUs and 4 GPUs each
- 4. NVLink switch tray · 9 × 1RU
- 5. NVLink switch chip · 2 per switch tray
- 6. Grace CPU · 2 per compute tray, 36 per rack
- 7. Blackwell GPU · 4 per compute tray, 72 per rack
- 8. ConnectX-7 NIC, 400 Gb/s · 4 per compute tray
- 9. Liquid-cooling manifold · rear, supply and return
- 10. Bus bar · rear, fed by the power shelves
Point at a part or a row. The other one lights.
GB200 NVL72 specifications.
| Parameter | Value | Scope | Source |
|---|---|---|---|
| Memory | |||
| Memory | 186 GB HBM3E | Per GPU | [1] |
| Memory bandwidth | 8 TB/s | Per GPU | [1] |
| GPU memory in the rack | 13.4 TB HBM3E | Per rack | [3] |
| GPU memory bandwidth | 576 TB/s | Per rack | [3] |
| CPU memory | 17 TB LPDDR5X · 14 TB/s | Per rack | [3] |
| Fast memory | Up to 31 TB | Per rack | [3] |
| Compute, peak · dense / sparse | |||
| FP4 Tensor Core | 10 / 20 PFLOPS | Per GPU | [1] |
| FP8 Tensor Core | 5 / 10 PFLOPS | Per GPU | [1] |
| BF16 Tensor Core | Not published | Per GPU | |
| FP4 Tensor Core | 720 / 1,440 PFLOPS | Per rack | [3] |
| FP8 Tensor Core | 360 / 720 PFLOPS | Per rack | [3] |
| Interconnect | |||
| NVLink 5 | 1.8 TB/s | Per GPU | [2] |
| NVLink domain | 72 GPUs in one NVLink domain · 130 TB/s across the rack | Per rack | [3] |
| Host interface | PCIe Gen5 · 128 GB/s | Per GPU | [1] |
| Scale-out network | 4 × ConnectX-7 400G per compute tray (NVIDIA DGX reference) | Per compute tray | [5] |
| Rack | |||
| GPUs | 72 | Per rack | [4] |
| Grace CPUs | 36 | Per rack | [4] |
| Compute trays | 18 · 2 Grace CPUs and 4 GPUs each | Per rack | [4] |
| NVLink switch trays | 9 | Per rack | [4] |
| Power and form factor | |||
| Power | Up to 1,200 W, configurable | Per GPU | [1] |
| Form factor | Liquid-cooled rack; Superchip modules (1 Grace CPU, 2 GPUs) | [1] | |
| Cooling | Liquid-cooled GPUs and CPUs; air for NICs and storage | Per rack | [4] |
| Rack power | About 120 kW (NVIDIA DGX GB reference) | Per rack | [4] |
Specifications are NVIDIA reference figures.
Best for
- Real-time trillion-parameter inference
- Large-scale training
- Data processing
Notes on the figures
- HBM total: 13.4 TB (product page; 72 × 186 GB). The datasheet figure of 13.5 TB is not used.
- Reserved capacity only. Not offered on demand or as spot, and not a Workspaces or Sandboxes shape.
Reserved capacity, quoted per rack.
Rack-scale systems are reserved capacity. Talk to us about term, delivery and configuration.
-
Reserved
Contact sales
Take-or-pay under a signed order form. No early cancellation without penalty (Terms §3).
-
Term and delivery
Set in the order form
Term, delivery and configuration are agreed for each rack before the order is signed.
-
On-demand and spot
Not offered
Reserved capacity never runs as spot.
GB200 NVL72 listing
- Status
- By request
- Runs as
- Reserved racks
- Capacity
- Reserved
- Tied to
- One GPU type, one region, one InfiniBand fabric
One rack, one NVLink domain.
18 compute trays, each with 2 Grace CPUs and 4 GPUs, and 9 NVLink switch trays join 72 GPUs in one NVLink domain.
One GB200 NVL72 rack holds 18 compute trays. Each compute tray holds 4 GPUs and 2 Grace CPUs, which makes 72 GPUs and 36 Grace CPUs. 9 NVLink switch trays join all 72 GPUs in one NVLink domain, 130 TB/s across the rack. The figure shows counts, not positions in the rack. Connections: Compute tray to NVLink switch tray: NVLink 5; NVLink switch tray to One NVLink domain: 72 GPUs; One NVLink domain to In the rack.
- GPUs
- 4
- Grace CPUs
- 2
18 trays4 × ConnectX-7 400G per compute tray
- per GPU
- 1.8 TB/s
9 trays- GPUs
- 72
- Grace CPUs
- 36
- NVLink 5
- 130 TB/s across the rack
- GPU memory
- 13.4 TB HBM3E
- GPU bandwidth
- 576 TB/s
- CPU memory
- 17 TB LPDDR5X · 14 TB/s
- Fast memory
- Up to 31 TB
- FP4
- 720 / 1,440 PFLOPS
- FP8
- 360 / 720 PFLOPS
Tensor Core peak, dense / sparse.
- cooling
- Liquid-cooled GPUs and CPUs; air for NICs
- and storage
- power
- About 120 kW (NVIDIA DGX GB reference)
- GPUs
- 4
- Grace CPUs
- 2
18 trays4 × ConnectX-7 400G per compute tray
- per GPU
- 1.8 TB/s
9 trays- GPUs
- 72
- Grace CPUs
- 36
- NVLink 5
- 130 TB/s across the rack
- GPU memory
- 13.4 TB HBM3E
- GPU bandwidth
- 576 TB/s
- CPU memory
- 17 TB LPDDR5X · 14 TB/s
- Fast memory
- Up to 31 TB
- FP4
- 720 / 1,440 PFLOPS
- FP8
- 360 / 720 PFLOPS
- cooling
- Liquid-cooled GPUs and CPUs; air
- for NICs and storage
- power
- About 120 kW (NVIDIA DGX GB
- reference)
- GPUs
- 4
- Grace CPUs
- 2
18 trays4 × ConnectX-7 400G per compute tray
- per GPU
- 1.8 TB/s
9 trays- GPUs
- 72
- Grace CPUs
- 36
- NVLink 5
- 130 TB/s across the rack
- GPU memory
- 13.4 TB HBM3E
- GPU bandwidth
- 576 TB/s
- CPU memory
- 17 TB LPDDR5X · 14 TB/s
- Fast memory
- Up to 31 TB
- FP4
- 720 / 1,440 PFLOPS
- FP8
- 360 / 720 PFLOPS
- cooling
- Liquid-cooled GPUs and CPUs;
- air for NICs and storage
- power
- About 120 kW (NVIDIA DGX GB
- reference)
NVIDIA GB200 NVL72 on Fantasti: fantasti.ai/gpus/gb200-nvl72 · Request capacity: fantasti.ai/contact
- Sheet
- 01 / 01
- Title
- GB200 NVL72 datasheet
- Figures
- NVIDIA reference
- Reviewed
- 2026-10-10
Request GB200 NVL72 capacity.
Rack-scale systems are reserved capacity. Talk to us about term, delivery and configuration.