Layer 2 · head to head
AMD Helios (Instinct MI455X) vs NVIDIA GB200 NVL72
432 GB per GPU of ROCm against the CUDA rack everyone already runs. The memory argument at rack scale.
| AMD Helios (Instinct MI455X) | NVIDIA GB200 NVL72 | |
|---|---|---|
| Can you get it | Rentable by the hour from multiple clouds; purchasable at scale. | Rentable by the hour from multiple clouds; purchasable at scale. |
| Memory | 31 TB of HBM4 per rack across 72 MI455X GPUs; 432 GB of HBM4 per GPU. | 13.5 TB HBM3E across the rack |
| Bandwidth | 260 TB/s scale-up per rack; 19.6 TB/s per GPU. | 576 TB/s aggregate |
| Compute | 2.9 exaFLOPS FP4 and 1.4 exaFLOPS FP8 per rack, on AMD's own figures. | 1.44 exaFLOPS FP4 inference |
| Power | Not published per rack. | ~120 kW nominal, 130-132 kW observed at full load |
| Interconnect | Scale-up fabric, Helios rack architecture | NVLink 5, 72 GPUs in one coherent domain |
| Software | ROCm | CUDA |
| Workload | both | both |
AMD Helios (Instinct MI455X)
For: Buyers with a working ROCm story who want more memory per GPU than NVIDIA sells, and a second source at rack scale.
The catch: ROCm is still the decision, not the silicon. The memory advantage per GPU is real and large — 432 GB against Blackwell-class parts — but it only pays if your stack runs on ROCm without a rewrite, and for most teams today it does not. Racks are quoted around $5-5.5M, which is one of very few public rack prices anywhere at this layer.
Economics: AMD's leverage here is memory capacity and being a second source, not price per FLOP. The commercial proof is commitment rather than benchmark: Oracle Cloud is the first hyperscaler offering a public MI450-series supercluster at 50,000 GPUs from Q3 2026, and OpenAI has committed to 6 GW of AMD Instinct with the first gigawatt in H2 2026.
NVIDIA GB200 NVL72
For: Frontier training and the largest inference deployments, where a single 72-GPU coherent memory domain is the point.
The catch: Liquid cooling is a mandatory architectural requirement, not a preference. At 120-132 kW a rack it renders most enterprise halls structurally and electrically unable to host it. This is a datacentre decision, not a hardware decision.
Economics: NVIDIA publishes two cents per million tokens and a 15x ROI claim for this configuration. Treat vendor TCO as a ceiling, not a forecast.
Before either — can you power it?
Not published per rack. against ~120 kW nominal, 130-132 kW observed at full load. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.