macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

Layer 2 · head to head

NVIDIA GB200 NVL72 vs Google TPU v7 (Ironwood)

Rack-scale coherent memory against a 9,216-chip inference pod. Different shapes of the same bet.

NVIDIA GB200 NVL72Google TPU v7 (Ironwood)
Can you get itRentable by the hour from multiple clouds; purchasable at scale.Cannot be bought. Exists only inside one cloud, so choosing it is choosing that cloud.
Memory13.5 TB HBM3E across the rack192 GB HBM3E per chip
Bandwidth576 TB/s aggregate7.37 TB/s per chip
Compute1.44 exaFLOPS FP4 inference4,614 FP8 TFLOPS per chip
Power~120 kW nominal, 130-132 kW observed at full loadNot published
InterconnectNVLink 5, 72 GPUs in one coherent domain9.6 Tb/s inter-chip (ICI)
SoftwareCUDAJAX, XLA, PyTorch/XLA
Workloadbothinference

NVIDIA GB200 NVL72

For: Frontier training and the largest inference deployments, where a single 72-GPU coherent memory domain is the point.

The catch: Liquid cooling is a mandatory architectural requirement, not a preference. At 120-132 kW a rack it renders most enterprise halls structurally and electrically unable to host it. This is a datacentre decision, not a hardware decision.

Economics: NVIDIA publishes two cents per million tokens and a 15x ROI claim for this configuration. Treat vendor TCO as a ceiling, not a forecast.

Google TPU v7 (Ironwood)

For: Inference at scale on Google Cloud. Ships in 256-chip and 9,216-chip configurations.

The catch: You cannot buy this. It exists only inside Google Cloud, so choosing it is choosing a cloud, permanently. It is also inference-optimised — do not benchmark it as a training part.

Economics: Vertically integrated: the price you see is a cloud price, not a chip price, and it is not comparable to a $/GPU-hour rental line.

Before either — can you power it?

~120 kW nominal, 130-132 kW observed at full load against Not published. In 2026 that comparison usually matters more than the FLOPS one: the US interconnection queue exceeds 2,600 GW with waits approaching five years, and roughly 80% of projects withdraw before energising.

Layer 1 — Energy · Nuclear vs gas · Direct-to-chip cooling

Specifications verified 2026-09-06. We take no commission at this layer, on either part.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.