Layer 2 · head to head
AMD Instinct MI350X vs AMD Instinct MI355X
Same 288 GB, same generation. The question is whether the flagship's extra compute changes your cost per token at all.
| AMD Instinct MI350X | AMD Instinct MI355X | |
|---|---|---|
| Can you get it | Rentable by the hour from multiple clouds; purchasable at scale. | Rentable by the hour from multiple clouds; purchasable at scale. |
| Memory | 288 GB HBM3E | 288 GB HBM3E |
| Bandwidth | 8 TB/s | 8 TB/s |
| Compute | FP4/FP8; below MI355X | ~10 PFLOPS sparse FP4 |
| Power | Not consistently published | Not consistently published |
| Interconnect | Infinity Fabric | Infinity Fabric |
| Software | ROCm | ROCm |
| Workload | both | both |
AMD Instinct MI350X
For: The volume part of the MI350 generation, with the same 288 GB memory advantage.
The catch: Same ROCm question as the MI355X. Check that your serving stack — vLLM, SGLang, TensorRT-LLM — actually supports your model on ROCm before committing.
Economics: AMD claims 2.6x the inference throughput of an H100 on Llama 3.1 405B.
AMD Instinct MI355X
For: Memory-bound inference. 288 GB against Blackwell's 192 GB is the cleanest technical argument AMD has made in a decade.
The catch: ROCm, not CUDA. The gap has narrowed sharply but it is still the reason procurement teams say no, and porting cost is real.
Economics: AMD positions this as the best price-per-token part for inference at scale, and claims 2.6x H100 inference throughput on Llama 3.1 405B for the MI350X.
Before either — can you power it?
Not consistently published against Not consistently published. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.