</>macrostackBrowse all
Tool profile · AI Voice & Speech

Chatterbox

MIT-licensed voice cloning that beat ElevenLabs in a blind test.

92
sovereignty

Chatterbox is an open text-to-speech model under the MIT licence with zero-shot voice cloning from roughly five seconds of reference audio, emotion/intensity control, and support for around 23 languages. In a published blind listening comparison, participants preferred Chatterbox-Turbo about 65% of the time against ElevenLabs' roughly 25% — a result worth treating as one data point rather than a verdict, but a striking one for a freely licensed model.

OPEN SOURCEMITSELF-HOSTLOCAL-FIRST
LicenseMIT
PricingFree — MIT licensed weights and code. Hardware only; a GPU is recommended for comfortable real-time cloning.
Open sourceYes
Self-hostableYes
Local-first dataYes

What it does well

  • +MIT licence — the most permissive terms available here
  • +Zero-shot voice cloning from about five seconds of audio
  • +Emotion and intensity control, roughly 23 languages
  • +Won a published blind preference test against ElevenLabs

Where it falls short

  • Heavier than Kokoro — wants a real GPU for smooth cloning
  • Voice cloning carries consent and likeness obligations that are yours to manage
  • Younger project, smaller ecosystem and less documentation

Chatterbox as an alternative to

Where Chatterbox shows up in our comparisons, and how it ranked.

Chatterbox head-to-head

Straight comparisons against the tools people weigh it against.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.