macrostack
Best of · 8 tools ranked

The best Local & Sovereign AI

Every option ranked — open-source, self-hostable, and commercial — by our transparent Sovereignty Score, with honest trade-offs so you choose what fits you, not us.

Run large language models on your own hardware — private, offline-capable, no per-token API bill.

L4Models & tooling — layer 4 of the AI stack
  1. 1

    Ollama

    Top pickOpen source

    Run Llama, Mistral, Qwen and more with one command.

    Free / self-host (you pay only for your own hardware + power) · in our OpenAI API (ChatGPT) comparison →

    92
    sovereignty
    What is Ollama? →
  2. 2

    LocalAI

    Open source

    A drop-in, OpenAI-compatible API you host yourself.

    Free / self-host · in our OpenAI API (ChatGPT) comparison →

    90
    sovereignty
    What is LocalAI? →
  3. 3

    vLLM

    Open source

    High-throughput serving for production-grade local inference.

    Free / self-host · in our OpenAI API (ChatGPT) comparison →

    88
    sovereignty
    What is vLLM? →
  4. 4

    Mistral AI

    Open source

    Open-weight Apache-2.0 models with an EU-hosted API — the exit ramp stays open.

    Open weights free to self-host; API from $0.04/1M tokens (Ministral 3B), Large at $2/$6 per 1M · in our OpenAI API (ChatGPT) comparison →

    70
    sovereignty
    What is Mistral AI? →
  5. 5

    LM Studio

    Commercial

    A polished desktop GUI for running local models.

    Free desktop app · in our OpenAI API (ChatGPT) comparison →

    68
    sovereignty
    What is LM Studio? →
  6. 6

    Together AI

    Commercial

    One API for the whole open-model universe — Llama, DeepSeek, Qwen and more.

    Usage-based, ~$0.05–$9 per 1M tokens by model; intro credits for new accounts · in our OpenAI API (ChatGPT) comparison →

    46
    sovereignty
    What is Together AI? →
  7. 7

    Groq

    Commercial

    The speed king — open models at 500+ tokens/second on custom LPU chips.

    Most models under $1 per 1M tokens; Llama 3.3 70B $0.59/$0.79; batch −50% · in our OpenAI API (ChatGPT) comparison →

    42
    sovereignty
    What is Groq? →
  8. 8

    Anthropic Claude API

    Commercial

    The frontier-quality closed alternative — strongest at reasoning and code.

    Haiku $1/$5 · Sonnet $3/$15 · Opus $5/$25 per 1M tokens; batch −50%, caching −90% · in our OpenAI API (ChatGPT) comparison →

    34
    sovereignty
    What is Anthropic Claude API? →

Is it free?

The free tier, its real limits and where you start paying — read from each vendor's own pricing page and dated.

Run these yourself

What each one actually needs — real RAM, honest running cost, and the setup time nobody quotes.

  • Self-hosting Ollama 8 GB VRAM for a 7B model at usable speed; 24 GB for 30B-class
  • Self-hosting LocalAI 8 GB RAM for CPU inference; 8 GB VRAM for anything comfortable
  • Self-hosting vLLM 24 GB VRAM minimum for useful production serving
  • Self-hosting Mistral AI 24 GB of VRAM for Ministral 3 14B in FP8 (stated on its card). Mistral Small 3.2 (24B) needs ~55 GB in bf16, or fits a single RTX 4090 once quantised per the Small 3.1 card. Small 4 (119B MoE) and Large 3 (675B MoE) are multi-GPU models.
  • Self-hosting LM Studio 16 GB system RAM, and VRAM is what actually decides your model

Replacing a specific tool?

Head-to-head comparisons for each popular local & sovereign ai product.

Straight head-to-heads

Two local & sovereign ai tools, side by side — verified facts and a plain verdict.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.