Layer 2.5 · head to head
Modal vs Baseten
Scale-to-zero, and which one's cold start you can live with.
| Modal | Baseten | |
|---|---|---|
| Can you leave | Sticky. Leaving means rewriting against a different interface, not changing a hostname. | Sticky. Leaving means rewriting against a different interface, not changing a hostname. |
| Counterparty | Low financially; high in code. Modal's programming model is genuinely delightful and genuinely non-portable — that is the trade. | You inherit whoever they buy from, which is not always disclosed. Ask. |
| What it is | You never see a GPU. You send a request and pay per token or per second of execution. | You never see a GPU. You send a request and pay per token or per second of execution. |
| How you reach it | You supply code; scaling and idle time are the provider's problem. | You call an endpoint. There is no infrastructure to see, and no infrastructure to move. |
| Accelerators | H100, A100, L40S, T4, and others | H100, A100, L4 and others behind managed deployments |
| Regions | US, EU | US |
| Pricing model | Per second of GPU time, scale to zero, no idle charge. | Per minute of compute for dedicated deployments; autoscaling to zero. |
| Getting started | Self-serve, generous free tier historically. | Self-serve. |
| Capacity | Managed pool. | Managed on top of other people's capacity. |
| Ownership | Private. | Private. |
Modal
For: Python teams who want infrastructure to disappear and are happy writing to a framework to get it.
The catch: Your deployment becomes Modal-shaped. Migrating off is a rewrite, not a redeploy. Price that in on day one, not year two.
Economics: Per-second billing with no idle cost is the cheapest possible shape for bursty inference. It is the wrong shape for a job that runs continuously.
Baseten
For: Teams deploying their own model weights who want autoscaling without operating Kubernetes.
The catch: Scale-to-zero is the headline and cold starts are the cost. For interactive products the cold-start number matters more than the hourly rate.
Economics: Pay for what you serve. Excellent for spiky traffic, poor value at steady high utilisation where a reserved GPU is cheaper.
Neither table row is a price
Deliberately. Published on-demand rates at this layer move weekly, and essentially nobody signing a real contract pays them — every serious buyer pays less than every list figure either of these companies publishes. Quoting one here would date this page within a month.
The GPU rental price index carries dated, sourced figures instead, and the durable finding there is the spread: the identical H100 rents from roughly $1.38 to $12.29 an hour depending only on who you rent it from.
The layers underneath both
Whichever you pick is renting you chips in a building that needs power. In 2026 that is the constraint that binds: Microsoft has disclosed an Azure backlog it cannot fill for want of megawatts rather than accelerators, and the US interconnection queue exceeds 2,600 GW with roughly 80% of projects withdrawing before they energise.
Layer 2 — Silicon · Layer 1 — Energy · The interconnection queue