Layer 2 · head to head
Google TPU v6e (Trillium) vs Google TPU v7 (Ironwood)
One generation apart, and Google publishes far more detail about the newer one — which is itself a signal about how they want them compared.
| Google TPU v6e (Trillium) | Google TPU v7 (Ironwood) | |
|---|---|---|
| Can you get it | Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud. | Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud. |
| Memory | Not published in directly comparable terms | 192 GB HBM3E per chip |
| Bandwidth | Not published | 7.37 TB/s per chip |
| Compute | Not published in comparable FP8 terms | 4,614 FP8 TFLOPS per chip |
| Power | Not published | Not published |
| Interconnect | ICI | 9.6 Tb/s inter-chip (ICI) |
| Software | JAX, XLA, PyTorch/XLA | JAX, XLA, PyTorch/XLA |
| Workload | both | inference |
Google TPU v6e (Trillium)
For: The previous TPU generation, still widely provisioned on Google Cloud and cheaper per hour than Ironwood.
The catch: Same lock-in as any TPU: Google Cloud only. Google publishes far less per-chip detail for v6e than for v7, which is itself a signal about how they want it compared.
Economics: Priced as a cloud SKU. Compare against v7 on tokens delivered per dollar, not on any spec sheet.
Google TPU v7 (Ironwood)
For: Inference at scale on Google Cloud. Ships in 256-chip and 9,216-chip configurations.
The catch: You cannot buy this. It exists only inside Google Cloud, so choosing it is choosing a cloud, permanently. It is also inference-optimised — do not benchmark it as a training part.
Economics: Vertically integrated: the price you see is a cloud price, not a chip price, and it is not comparable to a $/GPU-hour rental line.
Before either — can you power it?
Not published against Not published. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.