macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score
Best of · 8 tools ranked

The best Model serving & inference

Every option ranked — open-source, self-hostable, and commercial — by our transparent Sovereignty Score, with honest trade-offs so you choose what fits you, not us.

Once you have a model, something has to run it and answer requests. Serverless platforms rent the GPU by the second; the open engines let you own the endpoint.

L3Infrastructure — layer 3 of the AI stack
  1. 1

    vLLM

    Top pickOpen source

    The standard open inference engine — the thing under most serving platforms.

    Free and open source. You pay only for the GPUs you rent or own. · in our Modal comparison →

    95
    sovereignty
    What is vLLM? →
  2. 2

    Ollama

    Open source

    One command to a running model. The easiest way to stop paying per token.

    Free. Runs on hardware you already have. · in our Replicate comparison →

    95
    sovereignty
    What is Ollama? →
  3. 3

    LocalAI

    Open source

    A drop-in OpenAI replacement for chat, embeddings, images and audio.

    Free. No account, no telemetry, no usage cap. · in our Replicate comparison →

    94
    sovereignty
    What is LocalAI? →
  4. 4

    SGLang

    Open source

    Structured generation and prefix caching — the fast one for complex prompts.

    Free and unlimited; hardware costs are yours. · in our Replicate comparison →

    91
    sovereignty
    What is SGLang? →
  5. 5

    BentoML

    Open source

    Package any model as a container and deploy it wherever you like.

    Free and open source. BentoCloud is an optional paid hosted tier. · in our Modal comparison →

    90
    sovereignty
    What is BentoML? →
  6. 6

    Hugging Face TGI

    Open source

    Text Generation Inference — the production-hardened Rust serving stack.

    Free to self-host. Hugging Face sells a managed version if you want one. · in our Replicate comparison →

    90
    sovereignty
    What is Hugging Face TGI? →
  7. 7

    Ray Serve

    Open source

    Multi-node, multi-model serving for when one GPU is not the problem.

    Free and open source. Anyscale sells a managed Ray platform. · in our Modal comparison →

    89
    sovereignty
    What is Ray Serve? →
  8. 8

    RunPod Serverless

    CommercialPartner

    Scale-to-zero like Modal, at close to raw GPU rental prices.

    Per-second billing with scale-to-zero. H100 around $1.99/hr; no commitment. Rates observed 2026-07-30. · in our Modal comparison →

    56
    sovereignty
    What is RunPod Serverless? →

Run these yourself

What each one actually needs — real RAM, honest running cost, and the setup time nobody quotes.

Replacing a specific tool?

Head-to-head comparisons for each popular model serving & inference product.

Straight head-to-heads

Two model serving & inference tools, side by side — verified facts and a plain verdict.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.