macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score

Is Modal free? What it actually costs in 2026

The short answer

No free tier. It is a paid product.

Per-second billing with no minimum and scale-to-zero when idle. H100 capacity runs about $3.95/hr, A100 80GB about $2.50/hr, and L4 about $0.80/hr, with CPU and memory billed separately. For comparison, Replicate's H100 is around $5.49/hr and Baseten's around $6.50/hr per replica-hour billed continuously, while raw rental at RunPod is near $1.99/hr. Rates checked 2026-07-30; GPU pricing moves quickly, so confirm before planning around it.

When paying is still the right call

Stay for bursty or unpredictable traffic — scale-to-zero means an idle endpoint costs nothing, and no self-hosted setup matches that. Stay while you are still finding out whether the model is worth serving at all; paying twice the GPU rate for a week of iteration is trivial next to the time it saves. Stay if you have no Kubernetes and no desire to acquire it: vLLM on your own cluster is cheaper per hour and considerably more expensive in attention. And stay if your workload is genuinely spiky batch work, which is the shape serverless was built for.

What you are locked into

Higher than it looks, and it is worth naming precisely. The model weights and inference code are portable — that part is just Python and PyTorch. What is not portable is Modal's own layer: `@app.function`, image definitions, volumes, secrets and scheduling all live inside your source files, so leaving means unpicking the platform from the application rather than redeploying it elsewhere. Teams that keep the model logic in a plain module and confine Modal decorators to a thin entry point keep the exit cheap. Teams that do not, do not.

If you would rather not pay: vLLM

95

vLLM

The standard open inference engine — the thing under most serving platforms.

Free and open source. You pay only for the GPUs you rent or own.

See all 4 Modal alternatives compared

Compare the free options head-to-head

Pricing is verified against the vendor's published figures and dated above. Vendors change prices — confirm current numbers with Modal Labs before you commit. More model serving & inference decisions on Macrostack.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.