Google · TPU · Layer 2
Google TPU v7 (Ironwood)
Inference at scale on Google Cloud. Ships in 256-chip and 9,216-chip configurations.
cloud only Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
Specifications
- Memory
- 192 GB HBM3E per chip
- Bandwidth
- 7.37 TB/s per chip
- Compute
- 4,614 FP8 TFLOPS per chip
- Power
- Not published
- Interconnect
- 9.6 Tb/s inter-chip (ICI)
- Software
- JAX, XLA, PyTorch/XLA
- Form factor
- accelerator
- Workload
- inference
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
You cannot buy this. It exists only inside Google Cloud, so choosing it is choosing a cloud, permanently. It is also inference-optimised — do not benchmark it as a training part.
The economics
Vertically integrated: the price you see is a cloud price, not a chip price, and it is not comparable to a $/GPU-hour rental line.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- Google TPU v7 (Ironwood) vs AWS Trainium 3 — Two chips you cannot buy. This is a cloud commitment dressed as a hardware comparison.Compare with AWS Trainium 3
- NVIDIA GB200 NVL72 vs Google TPU v7 (Ironwood) — Rack-scale coherent memory against a 9,216-chip inference pod. Different shapes of the same bet.Compare with NVIDIA GB200 NVL72
- Google TPU v6e (Trillium) vs Google TPU v7 (Ironwood) — One generation apart, and Google publishes far more detail about the newer one — which is itself a signal about how they want them compared.Compare with Google TPU v6e (Trillium)
Related silicon
- Google TPU v6e (Trillium) — Not published in directly comparable terms, cloud only
- AWS Inferentia 2 — Not published in directly comparable terms, cloud only
- AWS Trainium 2 — Not published in directly comparable terms, cloud only
- AWS Trainium 3 — 144 GB HBM3e per chip, cloud only
- Cerebras WSE-3 — 44 GB on-chip SRAM, rent
- Groq LPU — 230 MB SRAM per chip, no HBM, rent