macrostack

Layer 5 · self-hosting reality check

What it actually takes to self-host Cline

The docs say Not stated for the extension itself, which runs inside VS Code or JetBrains. For local models, Cline's docs start at 16–32 GB of RAM for small or quantized models.. In practice you want 32 GB for a local model that actually works. Cline's tested minimum is Qwen3 Coder 30B at 4-bit (~17 GB download, 32K context, Compact Prompt on). 64 GB runs the 8-bit version (~32 GB) with 128K context and every feature. With an API key, any machine that runs VS Code is enough.. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.

Usually reached from Cursor alternatives, where Cline is one of the picks.

Wondering whether you need to at all? Is Cursor free? — what the free tier actually allows, and where the wall is.

RAM — documented minimumNot stated for the extension itself, which runs inside VS Code or JetBrains. For local models, Cline's docs start at 16–32 GB of RAM for small or quantized models.
RAM — what it really needs32 GB for a local model that actually works. Cline's tested minimum is Qwen3 Coder 30B at 4-bit (~17 GB download, 32K context, Compact Prompt on). 64 GB runs the 8-bit version (~32 GB) with 128K context and every feature. With an API key, any machine that runs VS Code is enough.
CPUAny modern CPU for the extension. Local inference needs a strong GPU or a machine with unified memory.
DiskAlmost nothing for the extension. Local models are 17–60 GB each (Qwen3 Coder 30B at 4-, 8- or 16-bit).
Monthly cost$0 — it runs on the machine you already own. Cline is free for individuals and you pay only for inference: per-token API use with your own key, or the hardware to run a local model.
Setup time5 minutes with an API key. An hour or two to set up Ollama or LM Studio with a big enough model and context.
How you install itInstall it from the VS Code or JetBrains marketplace (or `npm i -g cline` for the CLI), then pick a provider: your own API key, or Ollama / LM Studio serving a local model with Use Compact Prompt turned on.
Ongoing maintenanceAlmost none. The marketplace updates it automatically and releases come often (v4.1.22 on 30 Sep 2026). The real work is watching your API spend.
Where it stops scalingOne developer per install. Running locally, capability grows with memory: Cline's 128 GB+ tier (GLM-4.5-Air at 4-bit) is the one it describes as cloud-level performance.

The thing that catches people out

Cline's own testing found that models smaller than Qwen3 Coder 30B consistently fail with it: they produce broken output or refuse to run commands. That includes popular choices like gpt-oss-20b and an 8B DeepSeek R1 distill. Even the 30B model on a 32 GB machine only works with Compact Prompt on, which switches off MCP tools and Focus Chain. For local use, treat 32 GB of fast memory as the floor. Otherwise use an API key and set a spending limit at the provider before the agent runs on its own.

When not to self-host Cline

You expect a local model on a 16 GB laptop to replace a hosted frontier model: by Cline's own tests the small models fail. Skip it too if you can't accept per-token bills. An agent that reads files and calls tools in a loop spends much faster than chat, so use a capped API key or a cheaper hosted model.

Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.

Other Layer 5 self-hosting guides

Common questions

How much RAM does Cline actually need?
32 GB for a local model that actually works. Cline's tested minimum is Qwen3 Coder 30B at 4-bit (~17 GB download, 32K context, Compact Prompt on). 64 GB runs the 8-bit version (~32 GB) with 128K context and every feature. With an API key, any machine that runs VS Code is enough. in practice. The documented minimum is Not stated for the extension itself, which runs inside VS Code or JetBrains. For local models, Cline's docs start at 16–32 GB of RAM for small or quantized models., which is the figure at which the process starts rather than the figure at which it works under real use. Any modern CPU for the extension. Local inference needs a strong GPU or a machine with unified memory. alongside it.
What does self-hosting Cline cost per month?
$0 — it runs on the machine you already own. Cline is free for individuals and you pay only for inference: per-token API use with your own key, or the hardware to run a local model. This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for almost none. The marketplace updates it automatically and releases come often (v4.1.22 on 30 Sep 2026). The real work is watching your API spend.
How long does it take to set up Cline?
5 minutes with an API key. An hour or two to set up Ollama or LM Studio with a big enough model and context., via Install it from the VS Code or JetBrains marketplace (or `npm i -g cline` for the CLI), then pick a provider: your own API key, or Ollama / LM Studio serving a local model with Use Compact Prompt turned on..
When should I NOT self-host Cline?
You expect a local model on a 16 GB laptop to replace a hosted frontier model: by Cline's own tests the small models fail. Skip it too if you can't accept per-token bills. An agent that reads files and calls tools in a loop spends much faster than chat, so use a capped API key or a cheaper hosted model.
What is the most common mistake when self-hosting Cline?
Cline's own testing found that models smaller than Qwen3 Coder 30B consistently fail with it: they produce broken output or refuse to run commands. That includes popular choices like gpt-oss-20b and an 8B DeepSeek R1 distill. Even the 30B model on a 32 GB machine only works with Compact Prompt on, which switches off MCP tools and Focus Chain. For local use, treat 32 GB of fast memory as the floor. Otherwise use an API key and set a spending limit at the provider before the agent runs on its own.
The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.