Layer 2 · head to head
NVIDIA Vera Rubin VR200 NVL144 vs NVIDIA GB300 NVL72
One generation apart. The real question is not which is faster, it is whether you can get either.
| NVIDIA Vera Rubin VR200 NVL144 | NVIDIA GB300 NVL72 | |
|---|---|---|
| Can you get it | Rentable by the hour from multiple clouds; purchasable at scale. | Rentable by the hour from multiple clouds; purchasable at scale. |
| Memory | HBM4. NVIDIA quotes 20.7 TB per rack at NVL72 scale; the NVL144 CPX configuration is quoted at 100 TB of fast memory per rack. | Higher HBM3E capacity than GB200; per-rack figure not consistently published |
| Bandwidth | Around 20 TB/s per package on initial shipments, against an original 22 TB/s target the memory suppliers missed. 1.6-1.7 PB/s aggregate per rack. | Not published per rack |
| Compute | 8 exaFLOPS of FP4-class AI per NVL144 CPX rack, on NVIDIA's own figure. | Not published in comparable FP4 terms |
| Power | Not published in comparable per-rack terms at time of writing. | 135-200 kW per rack |
| Interconnect | NVLink, Rubin generation | NVLink 5 |
| Software | CUDA | CUDA |
| Workload | both | both |
NVIDIA Vera Rubin VR200 NVL144
For: Buyers who already committed to NVLink rack-scale and are taking the next generation. In practice that is a short list of hyperscalers and the largest neoclouds.
The catch: Effectively allocated rather than sold. First shipments began July 2026 and roughly 5,000-7,000 racks are expected across H2 2026 — a number worth weighing against the $725B of hyperscaler capex chasing them. The binding input is HBM4, which all three suppliers describe as committed or sold out through 2026, and NVIDIA already cut its bandwidth target because they could not hit it. A specification that moves before volume shipping is a supply story, not an engineering one.
Economics: The comparison that matters is against GB300, not against AMD: whether the throughput gain justifies waiting for an allocation you may not receive. For anyone outside the allocation list, the honest answer is that this part is not yet a purchasing option.
NVIDIA GB300 NVL72
For: The same buyers as GB200, one generation on, where throughput per megawatt is the binding constraint.
The catch: Pushes rack power to 135-200 kW. Very few facilities in the world can energise and cool that today, and the interconnection queue means new ones are a five-year decision.
Economics: NVIDIA claims up to 50x higher throughput per megawatt versus Hopper. Per-megawatt is the right frame in 2026 — see the Energy layer.
Before either — can you power it?
Not published in comparable per-rack terms at time of writing. against 135-200 kW per rack. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.