serverless inference · serverless · Layer 2.5
Modal
Python teams who want infrastructure to disappear and are happy writing to a framework to get it.
low portability Sticky. Leaving means rewriting against a different interface, not changing a hostname.
What you are buying
- Business model
- You never see a GPU. You send a request and pay per token or per second of execution.
- How you reach it
- You supply code; scaling and idle time are the provider's problem.
- Accelerators
- H100, A100, L40S, T4, and others
- Regions
- US, EU
- Pricing model
- Per second of GPU time, scale to zero, no idle charge.
- Getting started
- Self-serve, generous free tier historically.
- Capacity
- Managed pool.
- Ownership
- Private.
Verified 2026-09-09. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The counterparty
Low financially; high in code. Modal's programming model is genuinely delightful and genuinely non-portable — that is the trade.
A multi-year GPU commitment is a credit decision wearing a cloud contract. This is the section no benchmark covers and the one that decides what happens to your workload in 2028.
The catch
Your deployment becomes Modal-shaped. Migrating off is a rewrite, not a redeploy. Price that in on day one, not year two.
The economics
Per-second billing with no idle cost is the cheapest possible shape for bursty inference. It is the wrong shape for a job that runs continuously.
No rate is quoted on this page on purpose. Published list prices at this layer move weekly and essentially nobody signing a real contract pays them. Dated, sourced figures live in the GPU rental price index, where the spread between the cheapest and dearest seller of the identical chip runs to roughly 9x.
The layers underneath this one
Whatever you rent here is a chip in a building that needs power. In 2026 megawatts, not silicon, are the binding constraint on the whole industry — a frontier rack draws 120–200 kW against a 2026 average near 27 kW, and the US interconnection queue exceeds 2,600 GW.
Layer 2 — Silicon · Layer 1 — Energy · The interconnection queue · Tokens per watt
Compared against
- Modal vs Baseten — Scale-to-zero, and which one's cold start you can live with.Compare with Baseten
- Replicate vs Modal — Ship a feature in an afternoon, or own the deployment model.Compare with Replicate
- Fly.io vs Modal — Inference next to your app, or inference as its own platform.Compare with Fly.io
Related providers
- Baseten — serverless inference, low portability
- Fireworks AI — serverless inference, low portability
- Replicate — serverless inference, low portability
- Together AI — serverless inference, low portability
- Akash Network — marketplace, medium portability
- AWS (P5, P6, G6) — hyperscaler, medium portability