AMD · CDNA 3 · Layer 2
AMD Instinct MI300X
The first AMD part that was genuinely competitive for LLM inference, and now the value option in that lane.
rent Rentable by the hour from multiple clouds; purchasable at scale.
Specifications
- Memory
- 192 GB HBM3
- Bandwidth
- 5.3 TB/s
- Compute
- CDNA 3 FP8
- Power
- 750 W
- Interconnect
- Infinity Fabric
- Software
- ROCm
- Form factor
- accelerator
- Workload
- both
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
A generation behind MI355X. Worth choosing only on price, and only if ROCm already works for your stack.
The economics
For memory-bound LLM serving it achieves competitive or superior cost per token against H100 at most batch sizes.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- AMD Instinct MI300X vs NVIDIA H100 — The generation where AMD first became a real answer for memory-bound inference.Compare with NVIDIA H100
Related silicon
- AMD Instinct MI350X — 288 GB HBM3E, rent
- AMD Instinct MI355X — 288 GB HBM3E, rent
- NVIDIA B200 — 192 GB HBM3E, rent
- NVIDIA GB200 NVL72 — 13.5 TB HBM3E across the rack, rent
- NVIDIA GB300 NVL72 — Higher HBM3E capacity than GB200; per-rack figure not consistently published, rent
- NVIDIA H100 — 80 GB HBM3, rent