Layer 2.5 of the AI stack
Compute — who actually rents you the silicon
36 providers across 5 business models, rated on the axis nobody else publishes: who owns them, how the buildout is financed, and what your exit looks like if that goes wrong. Benchmarks tell you how fast a cluster is today. They do not tell you whether the debt behind it matures before your contract does.
Start with the spread
The same NVIDIA H100 SXM 80GB rents for $1.38 to $12.29 per GPU-hour — a 8.9x difference for a physically identical part. The chip is not the variable at this layer. The counterparty is.
The GPU rental price index → Surveyed 2026-09-09.
How to read this layer
- What are you actually buying?A neocloud sells you its own capacity. A marketplace brokers someone else’s. A serverless endpoint sells you an answer and never shows you a GPU. These are three different products at three different risks.
- Can you leave? Portability is the inverse of convenience here, and nobody selling the convenience puts that on the pricing page. 8 of these providers are a hostname change to exit; 5 are a rewrite.
- Who is the counterparty? Multi-year GPU commitments are credit decisions wearing a cloud contract. Ownership, financing and customer concentration are on every page below.
- What does list price actually mean? Mostly nothing. Every serious buyer at this layer pays less than every published rate, and the gap is the product.
What we do not rate, and who does
We do not benchmark clusters. SemiAnalysis’s ClusterMAX already rates these providers on hands-on testing — security, orchestration, storage, networking, reliability — across a far larger field than this one, by people who rent the hardware and run the tests. If you want to know how good a cluster is, read them.
This page answers the other half of the question, which nobody publishes: who you are signing with, and how you get out. Inventing benchmark numbers from a desk would make this page worse than useless, so we do not.
neocloud
Purpose-built GPU cloud. Owns or leases its own datacentre capacity and sells it directly.
- CoreWeave — H100, H200, GB200 NVL72, B200, A100, L40S. Portable. The workload is standard enough to move to another vendor without a rewrite.
- Nebius — H100, H200, B200, L40S. Movable with effort. Expect to redo images, storage wiring and networking.
- Lambda — H100, H200, B200, GH200, A100. Movable with effort. Expect to redo images, storage wiring and networking.
- Crusoe — H100, H200, B200, GB200. Movable with effort. Expect to redo images, storage wiring and networking.
- Nscale — H100, H200, B200, MI300X. Portable. The workload is standard enough to move to another vendor without a rewrite.
- Vultr — H100, H200, A100, L40S, MI300X, plus fractional GPUs. Movable with effort. Expect to redo images, storage wiring and networking.
- Civo — A100, L40S, and NVIDIA datacentre parts. Portable. The workload is standard enough to move to another vendor without a rewrite.
- Paperspace (DigitalOcean) — H100, A100, A6000, RTX 4000. Movable with effort. Expect to redo images, storage wiring and networking.
- Hyperstack (NexGen Cloud) — H100, H200, A100, L40S. Movable with effort. Expect to redo images, storage wiring and networking.
- DataCrunch — H100, H200, B200, A100, V100. Movable with effort. Expect to redo images, storage wiring and networking.
- GMI Cloud — H100, H200, B200, GB200. Portable. The workload is standard enough to move to another vendor without a rewrite.
- JarvisLabs — A100, H100, RTX 6000 Ada, A6000. Movable with effort. Expect to redo images, storage wiring and networking.
- Genesis Cloud — H100, HGX H100, RTX 3090/4090. Movable with effort. Expect to redo images, storage wiring and networking.
- Scaleway — H100, L40S, L4, plus Ascend and other non-NVIDIA options. Movable with effort. Expect to redo images, storage wiring and networking.
- OVHcloud — H100, A100, L40S, L4. Portable. The workload is standard enough to move to another vendor without a rewrite.
- Fly.io — A100, L40S. Movable with effort. Expect to redo images, storage wiring and networking.
- Voltage Park — H100 at scale. Portable. The workload is standard enough to move to another vendor without a rewrite.
marketplace
Brokers capacity somebody else owns. Cheapest headline rates, and the machine you get is not the machine you chose.
- RunPod — H100, H200, A100, L40S, RTX 4090, RTX 5090, and a long consumer tail. Movable with effort. Expect to redo images, storage wiring and networking.
- Vast.ai — Whatever the market is offering — RTX 5090, 4090, PRO 6000, V100, H100 and a long tail. Movable with effort. Expect to redo images, storage wiring and networking.
- Spheron — B200, H100, A100 and consumer parts. Movable with effort. Expect to redo images, storage wiring and networking.
- TensorDock — Broad consumer and datacentre mix. Movable with effort. Expect to redo images, storage wiring and networking.
- CUDO Compute — H100, A100, V100, consumer parts. Movable with effort. Expect to redo images, storage wiring and networking.
- Salad — Consumer GPUs only — RTX 3000/4000/5000 series. Movable with effort. Expect to redo images, storage wiring and networking.
- io.net — Aggregated consumer and datacentre GPUs. Movable with effort. Expect to redo images, storage wiring and networking.
- Akash Network — Aggregated, mostly consumer and older datacentre parts. Movable with effort. Expect to redo images, storage wiring and networking.
- SF Compute — H100 clusters. Portable. The workload is standard enough to move to another vendor without a rewrite.
serverless inference
You never see a GPU. You send a request and pay per token or per second of execution.
- Together AI — H100, H200, B200 behind the API; dedicated clusters available. Sticky. Leaving means rewriting against a different interface, not changing a hostname.
- Fireworks AI — Not disclosed per endpoint; NVIDIA datacentre class. Sticky. Leaving means rewriting against a different interface, not changing a hostname.
- Baseten — H100, A100, L4 and others behind managed deployments. Sticky. Leaving means rewriting against a different interface, not changing a hostname.
- Modal — H100, A100, L40S, T4, and others. Sticky. Leaving means rewriting against a different interface, not changing a hostname.
- Replicate — A100, H100, L40S, T4 behind hosted models. Sticky. Leaving means rewriting against a different interface, not changing a hostname.
hyperscaler
A general cloud that also rents accelerators. Most expensive per hour, and the one your compliance team has already approved.
- AWS (P5, P6, G6) — H100 (P5), B200 (P6), L4/L40S (G6), plus Trainium and Inferentia. Movable with effort. Expect to redo images, storage wiring and networking.
- Google Cloud (A3, A4) — H100, H200, B200, plus TPU v5e/v6/v7. Movable with effort. Expect to redo images, storage wiring and networking.
- Microsoft Azure (ND, NC) — H100, H200, GB200, A100, plus Maia internally. Movable with effort. Expect to redo images, storage wiring and networking.
- Oracle Cloud (OCI) — H100, H200, B200, GB200, A100. Portable. The workload is standard enough to move to another vendor without a rewrite.
aggregator
Sells access across other people's clouds through one contract and one API.
- Prime Intellect — Aggregated H100, A100 and others across providers. Movable with effort. Expect to redo images, storage wiring and networking.
How you reach the hardware
The abstraction you buy at decides your switching cost more than the logo does.
- bare metal
- You get the physical machine. Maximum control, and every layer above it is yours to run.
- kubernetes
- You get a cluster. Assumes you already run Kubernetes at scale.
- vm
- You get virtual machines. The familiar cloud model.
- container
- You push a container image and it runs. No cluster to operate.
- serverless
- You supply code; scaling and idle time are the provider's problem.
- api
- You call an endpoint. There is no infrastructure to see, and no infrastructure to move.
Head to head
30 pairs, authored rather than generated. 36 providers would produce 630 combinations, and that many near-identical tables is how a domain earns a thin-content judgement. Each of these is a comparison somebody actually makes.
- CoreWeave vs Lambda — Contracted scale against credit-card speed — and a list price roughly three times higher.
- CoreWeave vs Nebius — Two public companies at this layer, so for once you can read both sets of numbers before signing.
- Lambda vs RunPod — Real datacentre hourly against per-second marketplace billing.
- RunPod vs Vast.ai — Curated marketplace against open bidding: the price gap is the reliability gap.
- Vast.ai vs Salad — Two floors of the market — datacentre castoffs against gaming PCs.
- CoreWeave vs AWS (P5, P6, G6) — The 30-40% discount that is the neocloud pitch, against counterparty certainty.
- AWS (P5, P6, G6) vs Google Cloud (A3, A4) — Capacity Blocks against Dynamic Workload Scheduler — two answers to the same shortage.
- Google Cloud (A3, A4) vs Microsoft Azure (ND, NC) — TPU lock-in against enterprise-agreement pricing.
- AWS (P5, P6, G6) vs Oracle Cloud (OCI) — The most expensive hyperscaler against the cheapest, on egress as much as on GPUs.
- Together AI vs Fireworks AI — Per-token serving where the benchmark that matters is your prompt shape, not theirs.
- Modal vs Baseten — Scale-to-zero, and which one's cold start you can live with.
- Replicate vs Modal — Ship a feature in an afternoon, or own the deployment model.
- Together AI vs RunPod — Tokens against GPU-hours: the crossover calculation almost nobody does.
- Nebius vs Nscale — Two European builds, one public and one private, both selling sovereignty.
- Nscale vs DataCrunch — Nordic power economics at two very different scales.
- Crusoe vs Voltage Park — Two unusual ownership structures, two different reasons the price is low.
- Vultr vs Civo — Geographic spread against Kubernetes simplicity.
- OVHcloud vs Scaleway — European sovereignty from a public host against a telecoms-backed one.
- GMI Cloud vs CoreWeave — Roughly $2.60 against roughly $6.31 on an H200 — a $2,664 monthly gap per GPU at continuous run.
- Hyperstack (NexGen Cloud) vs DataCrunch — The middle band: cheaper than hyperscalers, more accountable than marketplaces.
- Spheron vs Vast.ai — Spot B200 access against the deepest consumer marketplace.
- Akash Network vs io.net — Two decentralised models, and what token incentives do to your supply.
- Prime Intellect vs CUDO Compute — Two aggregators, and the question both make harder to answer: whose datacentre is this?
- SF Compute vs Lambda — Buying a market window against buying a reservation.
- Paperspace (DigitalOcean) vs JarvisLabs — Notebook-first workflows at two very different scales.
- Fly.io vs Modal — Inference next to your app, or inference as its own platform.
- TensorDock vs RunPod — Two marketplaces, and how much curation is worth paying for.
- Genesis Cloud vs Nscale — Renewable Nordic compute at boutique and at industrial scale.
- Oracle Cloud (OCI) vs CoreWeave — Bare-metal RDMA clusters from a hyperscaler against a specialist.
- Microsoft Azure (ND, NC) vs RunPod — The widest price gap in the market — roughly $6.98 against roughly $2 for the same H100.
The layers either side of this one
Below: the accelerators themselves, and the power to run them. In 2026 megawatts are the binding constraint on the whole industry — Microsoft has disclosed an Azure order backlog it cannot fill because it cannot energise the capacity, not because it cannot buy the chips.
Layer 1 — Energy · Layer 2 — Silicon · Layer 3 — GPU compute tooling · The five-layer map