Groq · LPU · Layer 2
Groq LPU
Cheap, fast token generation via API. The lowest cost per million tokens among the specialist inference parts.
rent Rentable by the hour from multiple clouds; purchasable at scale.
Specifications
- Memory
- 230 MB SRAM per chip, no HBM
- Bandwidth
- 80 TB/s on-chip
- Compute
- Deterministic inference architecture
- Power
- Not published per chip
- Interconnect
- Chip-to-chip fabric
- Software
- GroqCloud API, GroqWare
- Form factor
- accelerator
- Workload
- inference
Verified 2026-09-06. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The catch
230 MB per chip means the model is split across many chips — Llama2-70B at Groq's headline speed required 576 chips, roughly eight racks. The per-chip spec is misleading unless you read it as a per-cluster architecture.
The economics
Cheaper per million tokens than Cerebras while being materially slower. If your product is priced per token rather than per millisecond, that is the right trade.
Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.
The power question underneath this
A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.
Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt
Compared against
- Cerebras WSE-3 vs Groq LPU — Speed against cost. Cerebras is 3-8x faster; Groq is cheaper per million tokens.Compare with Cerebras WSE-3
- Groq LPU vs NVIDIA H100 — Specialist inference silicon against the general-purpose default.Compare with NVIDIA H100
- SambaNova SN40L vs Groq LPU — Many models on one system against one model served cheaply.Compare with SambaNova SN40L
Related silicon
- Cerebras WSE-3 — 44 GB on-chip SRAM, rent
- SambaNova SN40L — 520 MiB SRAM + 64 GiB HBM + up to 1.5 TiB DDR per socket, rent
- AMD Instinct MI300X — 192 GB HBM3, rent
- AMD Instinct MI350X — 288 GB HBM3E, rent
- AMD Instinct MI355X — 288 GB HBM3E, rent
- AWS Inferentia 2 — Not published in directly comparable terms, cloud only