macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

NVIDIA · Vera Rubin · Layer 2

NVIDIA Vera Rubin VR200 NVL144

Buyers who already committed to NVLink rack-scale and are taking the next generation. In practice that is a short list of hyperscalers and the largest neoclouds.

rent Rentable by the hour from multiple clouds; purchasable at scale.

Specifications

Memory
HBM4. NVIDIA quotes 20.7 TB per rack at NVL72 scale; the NVL144 CPX configuration is quoted at 100 TB of fast memory per rack.
Bandwidth
Around 20 TB/s per package on initial shipments, against an original 22 TB/s target the memory suppliers missed. 1.6-1.7 PB/s aggregate per rack.
Compute
8 exaFLOPS of FP4-class AI per NVL144 CPX rack, on NVIDIA's own figure.
Power
Not published in comparable per-rack terms at time of writing.
Interconnect
NVLink, Rubin generation
Software
CUDA
Form factor
rack
Workload
both

Verified 2026-09-09. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.

The catch

Effectively allocated rather than sold. First shipments began July 2026 and roughly 5,000-7,000 racks are expected across H2 2026 — a number worth weighing against the $725B of hyperscaler capex chasing them. The binding input is HBM4, which all three suppliers describe as committed or sold out through 2026, and NVIDIA already cut its bandwidth target because they could not hit it. A specification that moves before volume shipping is a supply story, not an engineering one.

The economics

The comparison that matters is against GB300, not against AMD: whether the throughput gain justifies waiting for an allocation you may not receive. For anyone outside the allocation list, the honest answer is that this part is not yet a purchasing option.

Cost per million tokens is the metric that decides most real purchases, and there is no neutral benchmark for it — every published figure comes from a company selling one side of the comparison. Read all of them, ours included, as claims rather than measurements.

The power question underneath this

A rack of frontier accelerators draws 120–200 kW against a 2026 average of about 27 kW, and the US grid interconnection queue exceeds 2,600 GW with waits approaching five years. Whether you can energise this part is now a harder question than whether you can buy it.

Layer 1 — Energy · The interconnection queue · Rack power density · Tokens per watt

Where you actually rent this

Choosing the part is the easy half. The same accelerator rents for wildly different money depending only on who you rent it from — the H100 spread across published on-demand rates runs about 9x — and the provider, not the chip, is what decides whether you can leave later.

Layer 2.5 — Compute providers · The GPU rental price index

Compared against

Related silicon

Sources

Macrostack does not sell silicon and takes no commission at this layer — there is no affiliate programme for compute, which is exactly why nobody publishes a neutral comparison of it.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.