Layer 4 · self-hosting reality check
What it actually takes to self-host LiteLLM
The docs say 512 MB. In practice you want 2 GB. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.
Usually reached from Portkey alternatives, where LiteLLM is one of the picks.
Wondering whether you need to at all? Is Portkey free? — what the free tier actually allows, and where the wall is.
| RAM — documented minimum | 512 MB |
|---|---|
| RAM — what it really needs | 2 GB |
| CPU | 2 vCPU |
| Disk | Small |
| Monthly cost | $12–24/mo, against LLM gateway products charging per request |
| Setup time | 1 hour |
| How you install it | container; a YAML file maps model names to upstream providers |
| Ongoing maintenance | Low, but it is now in the path of every model call you make. |
| Where it stops scaling | Thousands of requests per second; it is a thin proxy. The upstream providers are the limit. |
The thing that catches people out
You have introduced a single point of failure between your application and every model provider. If the proxy is down, nothing reaches any model — including the fallback provider it exists to give you. Run at least two instances behind a load balancer, or keep a direct-to-provider code path you can flip to. A gateway that fails closed defeats the resilience it was added for.
When not to self-host LiteLLM
You call exactly one provider and have no plans to change. The abstraction earns its keep when you are routing, failing over or cost-tracking across several.
Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.
Other Layer 4 self-hosting guides
- Self-hosting Ollama8 GB VRAM for a 7B model at usable speed; 24 GB for 30B-class
- Self-hosting vLLM24 GB VRAM minimum for useful production serving
- Self-hosting QdrantVectors × dimensions × 4 bytes, in RAM, plus overhead — 1M × 768d is roughly 3 GB
- Self-hosting pgvector8 GB — the HNSW index wants to be resident
- Self-hosting LlamaIndex4 GB for the app; your vector store is the real cost
- Self-hosting faster-whisper5 GB VRAM for large-v3 in float16; 2 GB with int8
Common questions
- How much RAM does LiteLLM actually need?
- 2 GB in practice. The documented minimum is 512 MB, which is the figure at which the process starts rather than the figure at which it works under real use. 2 vCPU alongside it.
- What does self-hosting LiteLLM cost per month?
- $12–24/mo, against LLM gateway products charging per request This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for low, but it is now in the path of every model call you make.
- How long does it take to set up LiteLLM?
- 1 hour, via container; a YAML file maps model names to upstream providers.
- When should I NOT self-host LiteLLM?
- You call exactly one provider and have no plans to change. The abstraction earns its keep when you are routing, failing over or cost-tracking across several.
- What is the most common mistake when self-hosting LiteLLM?
- You have introduced a single point of failure between your application and every model provider. If the proxy is down, nothing reaches any model — including the fallback provider it exists to give you. Run at least two instances behind a load balancer, or keep a direct-to-provider code path you can flip to. A gateway that fails closed defeats the resilience it was added for.