macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

Layer 2 · head to head

SambaNova SN40L vs Groq LPU

Many models on one system against one model served cheaply.

SambaNova SN40LGroq LPU
Can you get itRentable by the hour from multiple clouds; purchasable at scale.Rentable by the hour from multiple clouds; purchasable at scale.
Memory520 MiB SRAM + 64 GiB HBM + up to 1.5 TiB DDR per socket230 MB SRAM per chip, no HBM
BandwidthOver 1 TB/s DDR-to-HBM model streaming80 TB/s on-chip
ComputeReconfigurable dataflowDeterministic inference architecture
PowerNot publishedNot published per chip
InterconnectSocket-to-socket fabricChip-to-chip fabric
SoftwareSambaFlowGroqCloud API, GroqWare
Workloadinferenceinference

SambaNova SN40L

For: Serving many large models from one system. The three-tier memory design exists to swap models in and out rather than to hold one resident.

The catch: The most unusual architecture here, so the least portable. Independently verified at 1,084 tokens/s for Llama3-8B — but on 16 sockets, which is the number that matters.

Economics: Its argument is models-per-dollar rather than tokens-per-dollar. If you serve one model, this is the wrong shape.

Groq LPU

For: Cheap, fast token generation via API. The lowest cost per million tokens among the specialist inference parts.

The catch: 230 MB per chip means the model is split across many chips — Llama2-70B at Groq's headline speed required 576 chips, roughly eight racks. The per-chip spec is misleading unless you read it as a per-cluster architecture.

Economics: Cheaper per million tokens than Cerebras while being materially slower. If your product is priced per token rather than per millisecond, that is the right trade.

Before either — can you power it?

Not published against Not published per chip. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.

Layer 1 — Energy · Nuclear vs gas · Direct-to-chip cooling

Specifications verified 2026-09-06. We take no commission at this layer, on either part.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.