macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

Layer 2 · head to head

Groq LPU vs NVIDIA H100

Specialist inference silicon against the general-purpose default.

Groq LPUNVIDIA H100
Can you get itRentable by the hour from multiple clouds; purchasable at scale.Rentable by the hour from multiple clouds; purchasable at scale.
Memory230 MB SRAM per chip, no HBM80 GB HBM3
Bandwidth80 TB/s on-chip3.35 TB/s
ComputeDeterministic inference architectureHopper-generation FP8
PowerNot published per chip700 W
InterconnectChip-to-chip fabricNVLink 4
SoftwareGroqCloud API, GroqWareCUDA
Workloadinferenceboth

Groq LPU

For: Cheap, fast token generation via API. The lowest cost per million tokens among the specialist inference parts.

The catch: 230 MB per chip means the model is split across many chips — Llama2-70B at Groq's headline speed required 576 chips, roughly eight racks. The per-chip spec is misleading unless you read it as a per-cluster architecture.

Economics: Cheaper per million tokens than Cerebras while being materially slower. If your product is priced per token rather than per millisecond, that is the right trade.

NVIDIA H100

For: The reference point everything else is benchmarked against, and still the most rentable accelerator on earth.

The catch: 80 GB is the binding limit. Large models need multi-GPU sharding that a 141 GB or 288 GB part would not, and sharding costs you latency and complexity.

Economics: The benchmark denominator. When a vendor claims '2.6x an H100', this is the H100 they mean.

Before either — can you power it?

Not published per chip against 700 W. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.

Layer 1 — Energy · Nuclear vs gas · Direct-to-chip cooling

Specifications verified 2026-09-06. We take no commission at this layer, on either part.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.