Chatterbox
MIT-licensed voice cloning that beat ElevenLabs in a blind test.
Chatterbox is an open text-to-speech model under the MIT licence with zero-shot voice cloning from roughly five seconds of reference audio, emotion/intensity control, and support for around 23 languages. In a published blind listening comparison, participants preferred Chatterbox-Turbo about 65% of the time against ElevenLabs' roughly 25% — a result worth treating as one data point rather than a verdict, but a striking one for a freely licensed model.
What it does well
- +MIT licence — the most permissive terms available here
- +Zero-shot voice cloning from about five seconds of audio
- +Emotion and intensity control, roughly 23 languages
- +Won a published blind preference test against ElevenLabs
Where it falls short
- −Heavier than Kokoro — wants a real GPU for smooth cloning
- −Voice cloning carries consent and likeness obligations that are yours to manage
- −Younger project, smaller ecosystem and less documentation
Chatterbox as an alternative to
Where Chatterbox shows up in our comparisons, and how it ranked.
Chatterbox head-to-head
Straight comparisons against the tools people weigh it against.