Amazon · Inferentia · Layer 2
AWS Inferentia 2
Cost-optimised inference on AWS for models the Neuron SDK supports well.
cloud only Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
Specifications
- Memory
- Not published in directly comparable terms
- Bandwidth
- Not published
- Compute
- Inference-optimised
- Power
- Not published
- Interconnect
- NeuronLink
- Software
- AWS Neuron SDK
- Form factor
- accelerator
- Workload
- inference
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
Model coverage is narrower than a GPU. Verify your exact model and quantisation compiles before you plan around the price.
The economics
The cheapest AWS inference silicon when your model is supported. Worth nothing at all when it is not.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- AWS Inferentia 2 vs AWS Trainium 3 — Inference-only and cheap against dense-training-capable. Model support decides this, not FLOPS.Compare with AWS Trainium 3
Related silicon
- AWS Trainium 2 — Not published in directly comparable terms, cloud only
- AWS Trainium 3 — 144 GB HBM3e per chip, cloud only
- Google TPU v7 (Ironwood) — 192 GB HBM3E per chip, cloud only
- Cerebras WSE-3 — 44 GB on-chip SRAM, rent
- Google TPU v6e (Trillium) — Not published in directly comparable terms, cloud only
- Groq LPU — 230 MB SRAM per chip, no HBM, rent