hyperscaler · vm · Layer 2.5
AWS (P5, P6, G6)
Anyone whose data, VPC, compliance boundary and team already live in AWS. The GPU price is rarely the deciding number.
medium portability Movable with effort. Expect to redo images, storage wiring and networking.
What you are buying
- Business model
- A general cloud that also rents accelerators. Most expensive per hour, and the one your compliance team has already approved.
- How you reach it
- You get virtual machines. The familiar cloud model.
- Accelerators
- H100 (P5), B200 (P6), L4/L40S (G6), plus Trainium and Inferentia
- Regions
- Global
- Pricing model
- On-demand, spot, Savings Plans, and Capacity Blocks for reserved GPU windows.
- Getting started
- Existing AWS account.
- Capacity
- Constrained on the newest parts; Capacity Blocks exist precisely because on-demand cannot be relied on.
- Ownership
- Amazon.
Verified 2026-09-09. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The counterparty
Effectively none. This is the reason the premium exists.
A multi-year GPU commitment is a credit decision wearing a cloud contract. This is the section no benchmark covers and the one that decides what happens to your workload in 2028.
The catch
Reported around $9.36/GPU-hour for a B200 Capacity Block against roughly $5.50 at Lambda. You are buying integration and counterparty certainty, and paying for both.
The economics
Egress and adjacency costs usually dominate the GPU line. Compare total workload cost, never the hourly rate alone.
No rate is quoted on this page on purpose. Published list prices at this layer move weekly and essentially nobody signing a real contract pays them. Dated, sourced figures live in the GPU rental price index, where the spread between the cheapest and dearest seller of the identical chip runs to roughly 9x.
The layers underneath this one
Whatever you rent here is a chip in a building that needs power. In 2026 megawatts, not silicon, are the binding constraint on the whole industry — a frontier rack draws 120–200 kW against a 2026 average near 27 kW, and the US interconnection queue exceeds 2,600 GW.
Layer 2 — Silicon · Layer 1 — Energy · The interconnection queue · Tokens per watt
Compared against
- CoreWeave vs AWS (P5, P6, G6) — The 30-40% discount that is the neocloud pitch, against counterparty certainty.Compare with CoreWeave
- AWS (P5, P6, G6) vs Google Cloud (A3, A4) — Capacity Blocks against Dynamic Workload Scheduler — two answers to the same shortage.Compare with Google Cloud (A3, A4)
- AWS (P5, P6, G6) vs Oracle Cloud (OCI) — The most expensive hyperscaler against the cheapest, on egress as much as on GPUs.Compare with Oracle Cloud (OCI)
Related providers
- Google Cloud (A3, A4) — hyperscaler, medium portability
- Microsoft Azure (ND, NC) — hyperscaler, medium portability
- Crusoe — neocloud, medium portability
- CUDO Compute — marketplace, medium portability
- DataCrunch — neocloud, medium portability
- Genesis Cloud — neocloud, medium portability