macrostack

Layer 5 · self-hosting reality check

What it actually takes to self-host Kokoro

The docs say Not stated by the project. The weights are 327 MB (kokoro-v1_0.pth) and the community ONNX export is 92 MB at 8-bit; it runs on CPU.. In practice you want Well under 2 GB of VRAM on a GPU — Kokoro-FastAPI measured 758 MiB freed on unload after a short clip and 1,656 MiB after 7.5 minutes of long-form audio on an RTX 4060 Ti. On CPU, an ordinary laptop with a couple of GB free is enough (estimate; not stated at source).. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.

Usually reached from ElevenLabs alternatives, where Kokoro is one of the picks.

Wondering whether you need to at all? Is ElevenLabs free? — what the free tier actually allows, and where the wall is.

RAM — documented minimumNot stated by the project. The weights are 327 MB (kokoro-v1_0.pth) and the community ONNX export is 92 MB at 8-bit; it runs on CPU.
RAM — what it really needsWell under 2 GB of VRAM on a GPU — Kokoro-FastAPI measured 758 MiB freed on unload after a short clip and 1,656 MiB after 7.5 minutes of long-form audio on an RTX 4060 Ti. On CPU, an ordinary laptop with a couple of GB free is enough (estimate; not stated at source).
CPUAny modern CPU: a third-party benchmark measured about 5x faster than realtime on a 32-vCPU AWS c6a, and Kokoro-FastAPI reports ~3.5 s to first audio on an older i7 and ~1 s on an M3 Pro. Any NVIDIA GPU takes it to 35–100x realtime.
DiskUnder 400 MB for the model: 327 MB of weights plus small voice files (the whole Hugging Face repo is 363 MB), plus the Python and PyTorch install
Monthly cost$0 — it runs on the machine you already own
Setup time15 minutes: pip install, install espeak-ng, generate a test clip
How you install itpip install kokoro soundfile, plus the espeak-ng system package (apt-get on Linux, the MSI installer on Windows); Kokoro-FastAPI wraps it in CPU or GPU Docker images with an OpenAI-compatible speech endpoint
Ongoing maintenanceLittle to maintain, and little arriving: the hexgrad/kokoro repo was last pushed on 2025-08-06 and the v1.0 weights date from January 2025. Pin kokoro, misaki and torch versions so a dependency bump cannot change pronunciation.
Where it stops scalingAt 35–100x realtime on one consumer GPU, a single instance renders an hour of audio in under two minutes; CPU handles one app or overnight batches. Kokoro-FastAPI adds streaming and an OpenAI-compatible endpoint when several apps share it.

The thing that catches people out

espeak-ng is not a pip dependency, and Kokoro does not fail without it. For English, the pipeline catches the missing library, logs "EspeakFallback not Enabled: OOD words will be skipped" and carries on — so names, brands and jargon that are not in misaki's dictionary silently vanish from the audio. Install espeak-ng before the first run, then test a sentence full of proper nouns and check the log for that warning.

When not to self-host Kokoro

You need a cloned or custom voice. Kokoro ships 54 preset voices in 8 languages, and its release has no encoder, so there is no way to make a new voice from a sample. For cloning, use Chatterbox (MIT) or stay on a hosted service like ElevenLabs.

Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.

Other Layer 5 self-hosting guides

Common questions

How much RAM does Kokoro actually need?
Well under 2 GB of VRAM on a GPU — Kokoro-FastAPI measured 758 MiB freed on unload after a short clip and 1,656 MiB after 7.5 minutes of long-form audio on an RTX 4060 Ti. On CPU, an ordinary laptop with a couple of GB free is enough (estimate; not stated at source). in practice. The documented minimum is Not stated by the project. The weights are 327 MB (kokoro-v1_0.pth) and the community ONNX export is 92 MB at 8-bit; it runs on CPU., which is the figure at which the process starts rather than the figure at which it works under real use. Any modern CPU: a third-party benchmark measured about 5x faster than realtime on a 32-vCPU AWS c6a, and Kokoro-FastAPI reports ~3.5 s to first audio on an older i7 and ~1 s on an M3 Pro. Any NVIDIA GPU takes it to 35–100x realtime. alongside it.
What does self-hosting Kokoro cost per month?
$0 — it runs on the machine you already own This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for little to maintain, and little arriving: the hexgrad/kokoro repo was last pushed on 2025-08-06 and the v1.0 weights date from January 2025. Pin kokoro, misaki and torch versions so a dependency bump cannot change pronunciation.
How long does it take to set up Kokoro?
15 minutes: pip install, install espeak-ng, generate a test clip, via pip install kokoro soundfile, plus the espeak-ng system package (apt-get on Linux, the MSI installer on Windows); Kokoro-FastAPI wraps it in CPU or GPU Docker images with an OpenAI-compatible speech endpoint.
When should I NOT self-host Kokoro?
You need a cloned or custom voice. Kokoro ships 54 preset voices in 8 languages, and its release has no encoder, so there is no way to make a new voice from a sample. For cloning, use Chatterbox (MIT) or stay on a hosted service like ElevenLabs.
What is the most common mistake when self-hosting Kokoro?
espeak-ng is not a pip dependency, and Kokoro does not fail without it. For English, the pipeline catches the missing library, logs "EspeakFallback not Enabled: OOD words will be skipped" and carries on — so names, brands and jargon that are not in misaki's dictionary silently vanish from the audio. Install espeak-ng before the first run, then test a sentence full of proper nouns and check the log for that warning.
The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.