macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

Layer 2.5 · head to head

Modal vs Baseten

Scale-to-zero, and which one's cold start you can live with.

ModalBaseten
Can you leaveSticky. Leaving means rewriting against a different interface, not changing a hostname.Sticky. Leaving means rewriting against a different interface, not changing a hostname.
CounterpartyLow financially; high in code. Modal's programming model is genuinely delightful and genuinely non-portable — that is the trade.You inherit whoever they buy from, which is not always disclosed. Ask.
What it isYou never see a GPU. You send a request and pay per token or per second of execution.You never see a GPU. You send a request and pay per token or per second of execution.
How you reach itYou supply code; scaling and idle time are the provider's problem.You call an endpoint. There is no infrastructure to see, and no infrastructure to move.
AcceleratorsH100, A100, L40S, T4, and othersH100, A100, L4 and others behind managed deployments
RegionsUS, EUUS
Pricing modelPer second of GPU time, scale to zero, no idle charge.Per minute of compute for dedicated deployments; autoscaling to zero.
Getting startedSelf-serve, generous free tier historically.Self-serve.
CapacityManaged pool.Managed on top of other people's capacity.
OwnershipPrivate.Private.

Modal

For: Python teams who want infrastructure to disappear and are happy writing to a framework to get it.

The catch: Your deployment becomes Modal-shaped. Migrating off is a rewrite, not a redeploy. Price that in on day one, not year two.

Economics: Per-second billing with no idle cost is the cheapest possible shape for bursty inference. It is the wrong shape for a job that runs continuously.

Baseten

For: Teams deploying their own model weights who want autoscaling without operating Kubernetes.

The catch: Scale-to-zero is the headline and cold starts are the cost. For interactive products the cold-start number matters more than the hourly rate.

Economics: Pay for what you serve. Excellent for spiky traffic, poor value at steady high utilisation where a reserved GPU is cheaper.

Neither table row is a price

Deliberately. Published on-demand rates at this layer move weekly, and essentially nobody signing a real contract pays them — every serious buyer pays less than every list figure either of these companies publishes. Quoting one here would date this page within a month.

The GPU rental price index carries dated, sourced figures instead, and the durable finding there is the spread: the identical H100 rents from roughly $1.38 to $12.29 an hour depending only on who you rent it from.

The layers underneath both

Whichever you pick is renting you chips in a building that needs power. In 2026 that is the constraint that binds: Microsoft has disclosed an Azure backlog it cannot fill for want of megawatts rather than accelerators, and the US interconnection queue exceeds 2,600 GW with roughly 80% of projects withdrawing before they energise.

Layer 2 — Silicon · Layer 1 — Energy · The interconnection queue

Verified 2026-09-09. We do not benchmark clusters and take no position paid for by either company. Where a provider here runs a referral programme it has not moved its placement — the comparison was written before any link was attached.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.