Kokoro
Top pick82M parameters, Apache-2.0, and the top of the TTS leaderboard.
Kokoro is a remarkably small open text-to-speech model — around 82 million parameters and a roughly 300MB footprint — that nonetheless ranks at the top of the TTS Arena leaderboard. It is released under Apache-2.0, so commercial use is free and unencumbered. Because it is small it runs faster than real time on modest hardware, including a cheap rented GPU, which makes it the most practical drop-in replacement for a per-character API for most teams.
What it does well
- +Apache-2.0 — commercial use with no restrictions
- +Tiny: ~300MB, runs faster than real time on modest hardware
- +Ranks top of the TTS Arena leaderboard despite its size
- +No per-character billing — cost stops scaling with usage
Where it falls short
- −No built-in voice cloning — you use the voices it ships with
- −Fewer languages than ElevenLabs' multilingual models
- −You own the deployment, updates, and uptime
Kokoro as an alternative to
Where Kokoro shows up in our comparisons, and how it ranked.
Kokoro head-to-head
Straight comparisons against the tools people weigh it against.