Kokoro
82M parameters, Apache-2.0, and the top of the TTS leaderboard.
Kokoro is a remarkably small open text-to-speech model — around 82 million parameters and a roughly 300MB footprint — that nonetheless ranks at the top of the TTS Arena leaderboard. It is released under Apache-2.0, so commercial use is free and unencumbered. Because it is small it runs faster than real time on modest hardware, including a cheap rented GPU, which makes it the most practical drop-in replacement for a per-character API for most teams.
Strengths
- +Apache-2.0 — commercial use with no restrictions
- +Tiny: ~300MB, runs faster than real time on modest hardware
- +Ranks top of the TTS Arena leaderboard despite its size
- +No per-character billing — cost stops scaling with usage
Trade-offs
- −No built-in voice cloning — you use the voices it ships with
- −Fewer languages than ElevenLabs' multilingual models
- −You own the deployment, updates, and uptime