Layer 2 · head to head
NVIDIA GB200 NVL72 vs Google TPU v7 (Ironwood)
Rack-scale coherent memory against a 9,216-chip inference pod. Different shapes of the same bet.
| NVIDIA GB200 NVL72 | Google TPU v7 (Ironwood) | |
|---|---|---|
| Can you get it | Rentable by the hour from multiple clouds; purchasable at scale. | Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud. |
| Memory | 13.5 TB HBM3E across the rack | 192 GB HBM3E per chip |
| Bandwidth | 576 TB/s aggregate | 7.37 TB/s per chip |
| Compute | 1.44 exaFLOPS FP4 inference | 4,614 FP8 TFLOPS per chip |
| Power | ~120 kW nominal, 130-132 kW observed at full load | Not published |
| Interconnect | NVLink 5, 72 GPUs in one coherent domain | 9.6 Tb/s inter-chip (ICI) |
| Software | CUDA | JAX, XLA, PyTorch/XLA |
| Workload | both | inference |
NVIDIA GB200 NVL72
For: Frontier training and the largest inference deployments, where a single 72-GPU coherent memory domain is the point.
The catch: Liquid cooling is a mandatory architectural requirement, not a preference. At 120-132 kW a rack it renders most enterprise halls structurally and electrically unable to host it. This is a datacentre decision, not a hardware decision.
Economics: NVIDIA publishes two cents per million tokens and a 15x ROI claim for this configuration. Treat vendor TCO as a ceiling, not a forecast.
Google TPU v7 (Ironwood)
For: Inference at scale on Google Cloud. Ships in 256-chip and 9,216-chip configurations.
The catch: You cannot buy this. It exists only inside Google Cloud, so choosing it is choosing a cloud, permanently. It is also inference-optimised — do not benchmark it as a training part.
Economics: Vertically integrated: the price you see is a cloud price, not a chip price, and it is not comparable to a $/GPU-hour rental line.
Before either — can you power it?
~120 kW nominal, 130-132 kW observed at full load against Not published. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.