AMD · CDNA 4 · Layer 2
AMD Instinct MI355X
Memory-bound inference. 288 GB against Blackwell's 192 GB is the cleanest technical argument AMD has made in a decade.
rent Rentable by the hour from multiple clouds; purchasable at scale.
Specifications
- Memory
- 288 GB HBM3E
- Bandwidth
- 8 TB/s
- Compute
- ~10 PFLOPS sparse FP4
- Power
- Not consistently published
- Interconnect
- Infinity Fabric
- Software
- ROCm
- Form factor
- accelerator
- Workload
- both
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
ROCm, not CUDA. The gap has narrowed sharply but it is still the reason procurement teams say no, and porting cost is real.
The economics
AMD positions this as the best price-per-token part for inference at scale, and claims 2.6x H100 inference throughput on Llama 3.1 405B for the MI350X.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- NVIDIA B200 vs AMD Instinct MI355X — The memory argument, stated cleanly: 192 GB of CUDA against 288 GB of ROCm.Compare with NVIDIA B200
- AWS Trainium 3 vs AMD Instinct MI355X — The two most credible cost-per-token challengers, one locked to a cloud, one not.Compare with AWS Trainium 3
- AMD Instinct MI350X vs AMD Instinct MI355X — Same 288 GB, same generation. The question is whether the flagship's extra compute changes your cost per token at all.Compare with AMD Instinct MI350X
Related silicon
- AMD Instinct MI300X — 192 GB HBM3, rent
- AMD Instinct MI350X — 288 GB HBM3E, rent
- NVIDIA B200 — 192 GB HBM3E, rent
- NVIDIA GB200 NVL72 — 13.5 TB HBM3E across the rack, rent
- NVIDIA GB300 NVL72 — Higher HBM3E capacity than GB200; per-rack figure not consistently published, rent
- NVIDIA H100 — 80 GB HBM3, rent