NVIDIA · Vera Rubin · Layer 2
NVIDIA Vera Rubin VR200 NVL144
Buyers who already committed to NVLink rack-scale and are taking the next generation. In practice that is a short list of hyperscalers and the largest neoclouds.
rent Rentable by the hour from multiple clouds; purchasable at scale.
Specifications
- Memory
- HBM4. NVIDIA quotes 20.7 TB per rack at NVL72 scale; the NVL144 CPX configuration is quoted at 100 TB of fast memory per rack.
- Bandwidth
- Around 20 TB/s per package on initial shipments, against an original 22 TB/s target the memory suppliers missed. 1.6-1.7 PB/s aggregate per rack.
- Compute
- 8 exaFLOPS of FP4-class AI per NVL144 CPX rack, on NVIDIA's own figure.
- Power
- Not published in comparable per-rack terms at time of writing.
- Interconnect
- NVLink, Rubin generation
- Software
- CUDA
- Form factor
- rack
- Workload
- both
Verified 2026-09-09. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
Effectively allocated rather than sold. First shipments began July 2026 and roughly 5,000-7,000 racks are expected across H2 2026 — a number worth weighing against the $725B of hyperscaler capex chasing them. The binding input is HBM4, which all three suppliers describe as committed or sold out through 2026, and NVIDIA already cut its bandwidth target because they could not hit it. A specification that moves before volume shipping is a supply story, not an engineering one.
The economics
The comparison that matters is against GB300, not against AMD: whether the throughput gain justifies waiting for an allocation you may not receive. For anyone outside the allocation list, the honest answer is that this part is not yet a purchasing option.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Where you actually rent this
Choosing the part is the easy half. The same accelerator rents for wildly different money depending only on who you rent it from — the H100 spread across published on-demand rates runs about 9x — and the provider, not the chip, is what decides whether you can leave later.
Compared against
- NVIDIA Vera Rubin VR200 NVL144 vs AMD Helios (Instinct MI455X) — The 2026 rack-scale fight. Both bet everything on HBM4, and HBM4 is sold out through the year.Compare with AMD Helios (Instinct MI455X)
- NVIDIA Vera Rubin VR200 NVL144 vs NVIDIA GB300 NVL72 — One generation apart. The real question is not which is faster, it is whether you can get either.Compare with NVIDIA GB300 NVL72
Related silicon
- NVIDIA B200 — 192 GB HBM3E, rent
- NVIDIA GB200 NVL72 — 13.5 TB HBM3E across the rack, rent
- NVIDIA GB300 NVL72 — Higher HBM3E capacity than GB200; per-rack figure not consistently published, rent
- NVIDIA H100 — 80 GB HBM3, rent
- NVIDIA H200 — 141 GB HBM3E, rent
- NVIDIA RTX PRO 6000 Blackwell — 96 GB GDDR7, buy