macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

Layer 2 · head to head

Cerebras WSE-3 vs Groq LPU

Speed against cost. Cerebras is 3-8x faster; Groq is cheaper per million tokens.

Cerebras WSE-3Groq LPU
Can you get itRentable by the hour from multiple clouds; purchasable at scale.Rentable by the hour from multiple clouds; purchasable at scale.
Memory44 GB on-chip SRAM230 MB SRAM per chip, no HBM
Bandwidth21 PB/s on-chip80 TB/s on-chip
Compute900,000 AI cores, 4 trillion transistorsDeterministic inference architecture
PowerNot published per unitNot published per chip
InterconnectOn-wafer fabricChip-to-chip fabric
SoftwareCerebras SDK, PyTorchGroqCloud API, GroqWare
Workloadinferenceinference

Cerebras WSE-3

For: Latency-critical inference. Delivered Llama 4 Maverick at 2,500 tokens per second per user — a figure no GPU cluster approaches.

The catch: 44 GB of SRAM is the whole memory system. There is no HBM to fall back on, so the model has to fit the architecture. This is a specialist part, not a general one.

Economics: Roughly 3-8x Groq's raw throughput depending on the benchmark, but Groq is cheaper per million tokens. Speed and cost point at different chips here.

Groq LPU

For: Cheap, fast token generation via API. The lowest cost per million tokens among the specialist inference parts.

The catch: 230 MB per chip means the model is split across many chips — Llama2-70B at Groq's headline speed required 576 chips, roughly eight racks. The per-chip spec is misleading unless you read it as a per-cluster architecture.

Economics: Cheaper per million tokens than Cerebras while being materially slower. If your product is priced per token rather than per millisecond, that is the right trade.

Before either — can you power it?

Not published per unit against Not published per chip. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.

Layer 1 — Energy · Nuclear vs gas · Direct-to-chip cooling

Specifications verified 2026-09-06. We take no commission at this layer, on either part.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.