Layer 2 · head to head
Google TPU v7 (Ironwood) vs AWS Trainium 3
Two chips you cannot buy. This is a cloud commitment dressed as a hardware comparison.
| Google TPU v7 (Ironwood) | AWS Trainium 3 | |
|---|---|---|
| Can you get it | Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud. | Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud. |
| Memory | 192 GB HBM3E per chip | 144 GB HBM3e per chip |
| Bandwidth | 7.37 TB/s per chip | 4.9 TB/s per chip |
| Compute | 4,614 FP8 TFLOPS per chip | 2.52 PFLOPS FP8 per chip |
| Power | Not published | Not published |
| Interconnect | 9.6 Tb/s inter-chip (ICI) | NeuronLink |
| Software | JAX, XLA, PyTorch/XLA | AWS Neuron SDK |
| Workload | inference | both |
Google TPU v7 (Ironwood)
For: Inference at scale on Google Cloud. Ships in 256-chip and 9,216-chip configurations.
The catch: You cannot buy this. It exists only inside Google Cloud, so choosing it is choosing a cloud, permanently. It is also inference-optimised — do not benchmark it as a training part.
Economics: Vertically integrated: the price you see is a cloud price, not a chip price, and it is not comparable to a $/GPU-hour rental line.
AWS Trainium 3
For: Dense training and inference on AWS. Unlike TPU v7 it is designed for both, which makes it the more honest Trainium-vs-TPU comparison.
The catch: AWS-only, and the Neuron SDK is a third software ecosystem to support alongside CUDA and ROCm. Model coverage is the question to ask, not FLOPS.
Economics: Independent analysis has singled out Trainium as the NVIDIA alternative that actually pays off on cost per token — largely because AWS prices it to move.
Before either — can you power it?
Not published against Not published. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.