Layer 5 · self-hosting reality check
What it actually takes to self-host Kokoro
The docs say Not stated by the project. The weights are 327 MB (kokoro-v1_0.pth) and the community ONNX export is 92 MB at 8-bit; it runs on CPU.. In practice you want Well under 2 GB of VRAM on a GPU — Kokoro-FastAPI measured 758 MiB freed on unload after a short clip and 1,656 MiB after 7.5 minutes of long-form audio on an RTX 4060 Ti. On CPU, an ordinary laptop with a couple of GB free is enough (estimate; not stated at source).. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.
Usually reached from ElevenLabs alternatives, where Kokoro is one of the picks.
Wondering whether you need to at all? Is ElevenLabs free? — what the free tier actually allows, and where the wall is.
| RAM — documented minimum | Not stated by the project. The weights are 327 MB (kokoro-v1_0.pth) and the community ONNX export is 92 MB at 8-bit; it runs on CPU. |
|---|---|
| RAM — what it really needs | Well under 2 GB of VRAM on a GPU — Kokoro-FastAPI measured 758 MiB freed on unload after a short clip and 1,656 MiB after 7.5 minutes of long-form audio on an RTX 4060 Ti. On CPU, an ordinary laptop with a couple of GB free is enough (estimate; not stated at source). |
| CPU | Any modern CPU: a third-party benchmark measured about 5x faster than realtime on a 32-vCPU AWS c6a, and Kokoro-FastAPI reports ~3.5 s to first audio on an older i7 and ~1 s on an M3 Pro. Any NVIDIA GPU takes it to 35–100x realtime. |
| Disk | Under 400 MB for the model: 327 MB of weights plus small voice files (the whole Hugging Face repo is 363 MB), plus the Python and PyTorch install |
| Monthly cost | $0 — it runs on the machine you already own |
| Setup time | 15 minutes: pip install, install espeak-ng, generate a test clip |
| How you install it | pip install kokoro soundfile, plus the espeak-ng system package (apt-get on Linux, the MSI installer on Windows); Kokoro-FastAPI wraps it in CPU or GPU Docker images with an OpenAI-compatible speech endpoint |
| Ongoing maintenance | Little to maintain, and little arriving: the hexgrad/kokoro repo was last pushed on 2025-08-06 and the v1.0 weights date from January 2025. Pin kokoro, misaki and torch versions so a dependency bump cannot change pronunciation. |
| Where it stops scaling | At 35–100x realtime on one consumer GPU, a single instance renders an hour of audio in under two minutes; CPU handles one app or overnight batches. Kokoro-FastAPI adds streaming and an OpenAI-compatible endpoint when several apps share it. |
The thing that catches people out
espeak-ng is not a pip dependency, and Kokoro does not fail without it. For English, the pipeline catches the missing library, logs "EspeakFallback not Enabled: OOD words will be skipped" and carries on — so names, brands and jargon that are not in misaki's dictionary silently vanish from the audio. Install espeak-ng before the first run, then test a sentence full of proper nouns and check the log for that warning.
When not to self-host Kokoro
You need a cloned or custom voice. Kokoro ships 54 preset voices in 8 languages, and its release has no encoder, so there is no way to make a new voice from a sample. For cloning, use Chatterbox (MIT) or stay on a hosted service like ElevenLabs.
Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.
Other Layer 5 self-hosting guides
- Self-hosting LibreOffice2 GB with a large spreadsheet open
- Self-hosting ONLYOFFICE6 GB for the Document Server with a handful of concurrent editors
- Self-hosting Collabora Online4 GB, and roughly 1 GB per 20 concurrent documents
- Self-hosting CryptPad2 GB for a small instance
- Self-hosting Mattermost4 GB for a team of 50 with PostgreSQL on the same box
- Self-hosting Rocket.Chat6 GB with MongoDB on the same machine
Common questions
- How much RAM does Kokoro actually need?
- Well under 2 GB of VRAM on a GPU — Kokoro-FastAPI measured 758 MiB freed on unload after a short clip and 1,656 MiB after 7.5 minutes of long-form audio on an RTX 4060 Ti. On CPU, an ordinary laptop with a couple of GB free is enough (estimate; not stated at source). in practice. The documented minimum is Not stated by the project. The weights are 327 MB (kokoro-v1_0.pth) and the community ONNX export is 92 MB at 8-bit; it runs on CPU., which is the figure at which the process starts rather than the figure at which it works under real use. Any modern CPU: a third-party benchmark measured about 5x faster than realtime on a 32-vCPU AWS c6a, and Kokoro-FastAPI reports ~3.5 s to first audio on an older i7 and ~1 s on an M3 Pro. Any NVIDIA GPU takes it to 35–100x realtime. alongside it.
- What does self-hosting Kokoro cost per month?
- $0 — it runs on the machine you already own This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for little to maintain, and little arriving: the hexgrad/kokoro repo was last pushed on 2025-08-06 and the v1.0 weights date from January 2025. Pin kokoro, misaki and torch versions so a dependency bump cannot change pronunciation.
- How long does it take to set up Kokoro?
- 15 minutes: pip install, install espeak-ng, generate a test clip, via pip install kokoro soundfile, plus the espeak-ng system package (apt-get on Linux, the MSI installer on Windows); Kokoro-FastAPI wraps it in CPU or GPU Docker images with an OpenAI-compatible speech endpoint.
- When should I NOT self-host Kokoro?
- You need a cloned or custom voice. Kokoro ships 54 preset voices in 8 languages, and its release has no encoder, so there is no way to make a new voice from a sample. For cloning, use Chatterbox (MIT) or stay on a hosted service like ElevenLabs.
- What is the most common mistake when self-hosting Kokoro?
- espeak-ng is not a pip dependency, and Kokoro does not fail without it. For English, the pipeline catches the missing library, logs "EspeakFallback not Enabled: OOD words will be skipped" and carries on — so names, brands and jargon that are not in misaki's dictionary silently vanish from the audio. Install espeak-ng before the first run, then test a sentence full of proper nouns and check the log for that warning.