Layer 5 · self-hosting reality check
What it actually takes to self-host Cline
The docs say Not stated for the extension itself, which runs inside VS Code or JetBrains. For local models, Cline's docs start at 16–32 GB of RAM for small or quantized models.. In practice you want 32 GB for a local model that actually works. Cline's tested minimum is Qwen3 Coder 30B at 4-bit (~17 GB download, 32K context, Compact Prompt on). 64 GB runs the 8-bit version (~32 GB) with 128K context and every feature. With an API key, any machine that runs VS Code is enough.. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.
Usually reached from Cursor alternatives, where Cline is one of the picks.
Wondering whether you need to at all? Is Cursor free? — what the free tier actually allows, and where the wall is.
| RAM — documented minimum | Not stated for the extension itself, which runs inside VS Code or JetBrains. For local models, Cline's docs start at 16–32 GB of RAM for small or quantized models. |
|---|---|
| RAM — what it really needs | 32 GB for a local model that actually works. Cline's tested minimum is Qwen3 Coder 30B at 4-bit (~17 GB download, 32K context, Compact Prompt on). 64 GB runs the 8-bit version (~32 GB) with 128K context and every feature. With an API key, any machine that runs VS Code is enough. |
| CPU | Any modern CPU for the extension. Local inference needs a strong GPU or a machine with unified memory. |
| Disk | Almost nothing for the extension. Local models are 17–60 GB each (Qwen3 Coder 30B at 4-, 8- or 16-bit). |
| Monthly cost | $0 — it runs on the machine you already own. Cline is free for individuals and you pay only for inference: per-token API use with your own key, or the hardware to run a local model. |
| Setup time | 5 minutes with an API key. An hour or two to set up Ollama or LM Studio with a big enough model and context. |
| How you install it | Install it from the VS Code or JetBrains marketplace (or `npm i -g cline` for the CLI), then pick a provider: your own API key, or Ollama / LM Studio serving a local model with Use Compact Prompt turned on. |
| Ongoing maintenance | Almost none. The marketplace updates it automatically and releases come often (v4.1.22 on 30 Sep 2026). The real work is watching your API spend. |
| Where it stops scaling | One developer per install. Running locally, capability grows with memory: Cline's 128 GB+ tier (GLM-4.5-Air at 4-bit) is the one it describes as cloud-level performance. |
The thing that catches people out
Cline's own testing found that models smaller than Qwen3 Coder 30B consistently fail with it: they produce broken output or refuse to run commands. That includes popular choices like gpt-oss-20b and an 8B DeepSeek R1 distill. Even the 30B model on a 32 GB machine only works with Compact Prompt on, which switches off MCP tools and Focus Chain. For local use, treat 32 GB of fast memory as the floor. Otherwise use an API key and set a spending limit at the provider before the agent runs on its own.
When not to self-host Cline
You expect a local model on a 16 GB laptop to replace a hosted frontier model: by Cline's own tests the small models fail. Skip it too if you can't accept per-token bills. An agent that reads files and calls tools in a loop spends much faster than chat, so use a capped API key or a cheaper hosted model.
Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.
Other Layer 5 self-hosting guides
- Self-hosting LibreOffice2 GB with a large spreadsheet open
- Self-hosting ONLYOFFICE6 GB for the Document Server with a handful of concurrent editors
- Self-hosting Collabora Online4 GB, and roughly 1 GB per 20 concurrent documents
- Self-hosting CryptPad2 GB for a small instance
- Self-hosting Mattermost4 GB for a team of 50 with PostgreSQL on the same box
- Self-hosting Rocket.Chat6 GB with MongoDB on the same machine
Common questions
- How much RAM does Cline actually need?
- 32 GB for a local model that actually works. Cline's tested minimum is Qwen3 Coder 30B at 4-bit (~17 GB download, 32K context, Compact Prompt on). 64 GB runs the 8-bit version (~32 GB) with 128K context and every feature. With an API key, any machine that runs VS Code is enough. in practice. The documented minimum is Not stated for the extension itself, which runs inside VS Code or JetBrains. For local models, Cline's docs start at 16–32 GB of RAM for small or quantized models., which is the figure at which the process starts rather than the figure at which it works under real use. Any modern CPU for the extension. Local inference needs a strong GPU or a machine with unified memory. alongside it.
- What does self-hosting Cline cost per month?
- $0 — it runs on the machine you already own. Cline is free for individuals and you pay only for inference: per-token API use with your own key, or the hardware to run a local model. This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for almost none. The marketplace updates it automatically and releases come often (v4.1.22 on 30 Sep 2026). The real work is watching your API spend.
- How long does it take to set up Cline?
- 5 minutes with an API key. An hour or two to set up Ollama or LM Studio with a big enough model and context., via Install it from the VS Code or JetBrains marketplace (or `npm i -g cline` for the CLI), then pick a provider: your own API key, or Ollama / LM Studio serving a local model with Use Compact Prompt turned on..
- When should I NOT self-host Cline?
- You expect a local model on a 16 GB laptop to replace a hosted frontier model: by Cline's own tests the small models fail. Skip it too if you can't accept per-token bills. An agent that reads files and calls tools in a loop spends much faster than chat, so use a capped API key or a cheaper hosted model.
- What is the most common mistake when self-hosting Cline?
- Cline's own testing found that models smaller than Qwen3 Coder 30B consistently fail with it: they produce broken output or refuse to run commands. That includes popular choices like gpt-oss-20b and an 8B DeepSeek R1 distill. Even the 30B model on a 32 GB machine only works with Compact Prompt on, which switches off MCP tools and Focus Chain. For local use, treat 32 GB of fast memory as the floor. Otherwise use an API key and set a spending limit at the provider before the agent runs on its own.