Layer 2 of the AI stack
Chips & silicon
24 accelerators from 13 vendors, with the field every other comparison leaves out: whether you can actually get one. Only 5 of these can be bought outright. 6 exist solely inside a single cloud, which makes choosing them a decision about a vendor rather than a chip.
How to read this layer
- Can you get it? Availability first, always. A part you cannot buy is not an option, however good the specification is.
- Does the model fit? Memory capacity decides that. Bandwidth decides how fast it runs once it does.
- What does useful work cost? Per token, per watt. There is no neutral benchmark for either — treat every published figure, including the ones quoted here, as a vendor claim.
- What software are you marrying? CUDA, ROCm, Neuron, XLA. This outlives the hardware and is the real switching cost.
NVIDIA
- NVIDIA GB200 NVL72 — 13.5 TB HBM3E across the rack, 576 TB/s aggregate. Rentable by the hour from multiple clouds; purchasable at scale.
- NVIDIA GB300 NVL72 — Higher HBM3E capacity than GB200; per-rack figure not consistently published, Not published per rack. Rentable by the hour from multiple clouds; purchasable at scale.
- NVIDIA B200 — 192 GB HBM3E, 8 TB/s. Rentable by the hour from multiple clouds; purchasable at scale.
- NVIDIA H200 — 141 GB HBM3E, 4.8 TB/s. Rentable by the hour from multiple clouds; purchasable at scale.
- NVIDIA H100 — 80 GB HBM3, 3.35 TB/s. Rentable by the hour from multiple clouds; purchasable at scale.
- NVIDIA RTX PRO 6000 Blackwell — 96 GB GDDR7, 1.8 TB/s. You can purchase this outright.
- NVIDIA DGX Spark — 128 GB unified LPDDR5X, 273 GB/s. You can purchase this outright.
Amazon
- AWS Trainium 3 — 144 GB HBM3e per chip, 4.9 TB/s per chip. Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
- AWS Trainium 2 — Not published in directly comparable terms, Not published. Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
- AWS Inferentia 2 — Not published in directly comparable terms, Not published. Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
AMD
- AMD Instinct MI355X — 288 GB HBM3E, 8 TB/s. Rentable by the hour from multiple clouds; purchasable at scale.
- AMD Instinct MI350X — 288 GB HBM3E, 8 TB/s. Rentable by the hour from multiple clouds; purchasable at scale.
- AMD Instinct MI300X — 192 GB HBM3, 5.3 TB/s. Rentable by the hour from multiple clouds; purchasable at scale.
- Google TPU v7 (Ironwood) — 192 GB HBM3E per chip, 7.37 TB/s per chip. Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
- Google TPU v6e (Trillium) — Not published in directly comparable terms, Not published. Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
Cerebras
- Cerebras WSE-3 — 44 GB on-chip SRAM, 21 PB/s on-chip. Rentable by the hour from multiple clouds; purchasable at scale.
Groq
- Groq LPU — 230 MB SRAM per chip, no HBM, 80 TB/s on-chip. Rentable by the hour from multiple clouds; purchasable at scale.
Huawei
- Huawei Ascend 910C — Not reliably published outside China, Not reliably published. Export controls or regional restrictions govern access.
Intel
- Intel Gaudi 3 — 128 GB HBM2E, 3.7 TB/s. You can purchase this outright.
Meta
- Meta MTIA v2 — Not published, Not published. Not available to anyone outside the company that built it.
Microsoft
- Microsoft Maia 200 — Not published, Not published. Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
Qualcomm
- Qualcomm Cloud AI 100 Ultra — 128 GB LPDDR, Below HBM parts. You can purchase this outright.
SambaNova
- SambaNova SN40L — 520 MiB SRAM + 64 GiB HBM + up to 1.5 TiB DDR per socket, Over 1 TB/s DDR-to-HBM model streaming. Rentable by the hour from multiple clouds; purchasable at scale.
Tenstorrent
- Tenstorrent Blackhole — GDDR6 rather than HBM, Below HBM parts by design. You can purchase this outright.
Head to head
21 pairs, authored rather than generated. Twenty-four chips would produce 276 combinations, and 276 near-identical spec tables is how a site earns a thin-content judgement. Each of these is a comparison somebody actually makes.
- NVIDIA B200 vs AMD Instinct MI355X — The memory argument, stated cleanly: 192 GB of CUDA against 288 GB of ROCm.
- Google TPU v7 (Ironwood) vs AWS Trainium 3 — Two chips you cannot buy. This is a cloud commitment dressed as a hardware comparison.
- NVIDIA GB200 NVL72 vs Google TPU v7 (Ironwood) — Rack-scale coherent memory against a 9,216-chip inference pod. Different shapes of the same bet.
- NVIDIA H100 vs NVIDIA H200 — The same architecture with 61 GB and 1.45 TB/s more. Whether that matters is decided entirely by your model size.
- AMD Instinct MI300X vs NVIDIA H100 — The generation where AMD first became a real answer for memory-bound inference.
- Intel Gaudi 3 vs NVIDIA H100 — Standard Ethernet and a lower price against the deepest software ecosystem in computing.
- Cerebras WSE-3 vs Groq LPU — Speed against cost. Cerebras is 3-8x faster; Groq is cheaper per million tokens.
- Groq LPU vs NVIDIA H100 — Specialist inference silicon against the general-purpose default.
- AWS Trainium 3 vs AMD Instinct MI355X — The two most credible cost-per-token challengers, one locked to a cloud, one not.
- NVIDIA GB300 NVL72 vs NVIDIA GB200 NVL72 — One generation, and 15-80 kW more per rack. The question is whether your facility can energise it.
- Huawei Ascend 910C vs NVIDIA H100 — A jurisdiction comparison, not a performance one.
- NVIDIA RTX PRO 6000 Blackwell vs NVIDIA DGX Spark — Both buyable, both local. Bandwidth against unified memory capacity.
- SambaNova SN40L vs Groq LPU — Many models on one system against one model served cheaply.
- Qualcomm Cloud AI 100 Ultra vs NVIDIA H200 — 150 W against 700 W. Tokens per watt is a different contest from tokens per second.
- Tenstorrent Blackhole vs NVIDIA RTX PRO 6000 Blackwell — The only fully open stack against the one everything is written for.
- AMD Instinct MI350X vs AMD Instinct MI355X — Same 288 GB, same generation. The question is whether the flagship's extra compute changes your cost per token at all.
- Google TPU v6e (Trillium) vs Google TPU v7 (Ironwood) — One generation apart, and Google publishes far more detail about the newer one — which is itself a signal about how they want them compared.
- AWS Trainium 2 vs AWS Trainium 3 — Region availability against raw capability. On AWS the older part is often the one you can actually get.
- AWS Inferentia 2 vs AWS Trainium 3 — Inference-only and cheap against dense-training-capable. Model support decides this, not FLOPS.
- Microsoft Maia 200 vs Meta MTIA v2 — Two chips nobody outside their owners can touch. They matter as the GPU orders that never get placed.
- Meta MTIA v2 vs NVIDIA B200 — The custom-silicon question in one line: every MTIA Meta deploys is a B200 NVIDIA does not sell.
The layer underneath
None of this runs without power, and in 2026 power is the harder problem. A single frontier rack draws 120–200 kW; the US grid interconnection queue exceeds 2,600 GW with waits approaching five years and an 80% withdrawal rate.