NVIDIA · Blackwell · Layer 2
NVIDIA GB200 NVL72
Frontier training and the largest inference deployments, where a single 72-GPU coherent memory domain is the point.
rent Rentable by the hour from multiple clouds; purchasable at scale.
Specifications
- Memory
- 13.5 TB HBM3E across the rack
- Bandwidth
- 576 TB/s aggregate
- Compute
- 1.44 exaFLOPS FP4 inference
- Power
- ~120 kW nominal, 130-132 kW observed at full load
- Interconnect
- NVLink 5, 72 GPUs in one coherent domain
- Software
- CUDA
- Form factor
- rack
- Workload
- both
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
Liquid cooling is a mandatory architectural requirement, not a preference. At 120-132 kW a rack it renders most enterprise halls structurally and electrically unable to host it. This is a datacentre decision, not a hardware decision.
The economics
NVIDIA publishes two cents per million tokens and a 15x ROI claim for this configuration. Treat vendor TCO as a ceiling, not a forecast.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- NVIDIA GB200 NVL72 vs Google TPU v7 (Ironwood) — Rack-scale coherent memory against a 9,216-chip inference pod. Different shapes of the same bet.Compare with Google TPU v7 (Ironwood)
- NVIDIA GB300 NVL72 vs NVIDIA GB200 NVL72 — One generation, and 15-80 kW more per rack. The question is whether your facility can energise it.Compare with NVIDIA GB300 NVL72
Related silicon
- NVIDIA B200 — 192 GB HBM3E, rent
- NVIDIA GB300 NVL72 — Higher HBM3E capacity than GB200; per-rack figure not consistently published, rent
- NVIDIA H100 — 80 GB HBM3, rent
- NVIDIA H200 — 141 GB HBM3E, rent
- NVIDIA RTX PRO 6000 Blackwell — 96 GB GDDR7, buy
- AMD Instinct MI300X — 192 GB HBM3, rent