Amazon · Trainium · Layer 2
AWS Trainium 3
Dense training and inference on AWS. Unlike TPU v7 it is designed for both, which makes it the more honest Trainium-vs-TPU comparison.
cloud only Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
Specifications
- Memory
- 144 GB HBM3e per chip
- Bandwidth
- 4.9 TB/s per chip
- Compute
- 2.52 PFLOPS FP8 per chip
- Power
- Not published
- Interconnect
- NeuronLink
- Software
- AWS Neuron SDK
- Form factor
- accelerator
- Workload
- both
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
AWS-only, and the Neuron SDK is a third software ecosystem to support alongside CUDA and ROCm. Model coverage is the question to ask, not FLOPS.
The economics
Independent analysis has singled out Trainium as the NVIDIA alternative that actually pays off on cost per token — largely because AWS prices it to move.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- Google TPU v7 (Ironwood) vs AWS Trainium 3 — Two chips you cannot buy. This is a cloud commitment dressed as a hardware comparison.Compare with Google TPU v7 (Ironwood)
- AWS Trainium 3 vs AMD Instinct MI355X — The two most credible cost-per-token challengers, one locked to a cloud, one not.Compare with AMD Instinct MI355X
- AWS Trainium 2 vs AWS Trainium 3 — Region availability against raw capability. On AWS the older part is often the one you can actually get.Compare with AWS Trainium 2
- AWS Inferentia 2 vs AWS Trainium 3 — Inference-only and cheap against dense-training-capable. Model support decides this, not FLOPS.Compare with AWS Inferentia 2
Related silicon
- AWS Trainium 2 — Not published in directly comparable terms, cloud only
- AWS Inferentia 2 — Not published in directly comparable terms, cloud only
- Google TPU v6e (Trillium) — Not published in directly comparable terms, cloud only
- Microsoft Maia 200 — Not published, cloud only
- AMD Instinct MI300X — 192 GB HBM3, rent
- AMD Instinct MI350X — 288 GB HBM3E, rent