</>macrostackBrowse all
Tool profile · AI Voice & Speech

Kokoro

Top pick

82M parameters, Apache-2.0, and the top of the TTS leaderboard.

94
sovereignty

Kokoro is a remarkably small open text-to-speech model — around 82 million parameters and a roughly 300MB footprint — that nonetheless ranks at the top of the TTS Arena leaderboard. It is released under Apache-2.0, so commercial use is free and unencumbered. Because it is small it runs faster than real time on modest hardware, including a cheap rented GPU, which makes it the most practical drop-in replacement for a per-character API for most teams.

OPEN SOURCEApache-2.0SELF-HOSTLOCAL-FIRST
LicenseApache-2.0
PricingFree — weights are Apache-2.0. Costs are hardware only; a rented GPU capable of faster-than-real-time synthesis runs roughly $20–$80/month, and nothing at all if you already own a suitable card.
Open sourceYes
Self-hostableYes
Local-first dataYes

What it does well

  • +Apache-2.0 — commercial use with no restrictions
  • +Tiny: ~300MB, runs faster than real time on modest hardware
  • +Ranks top of the TTS Arena leaderboard despite its size
  • +No per-character billing — cost stops scaling with usage

Where it falls short

  • No built-in voice cloning — you use the voices it ships with
  • Fewer languages than ElevenLabs' multilingual models
  • You own the deployment, updates, and uptime

Kokoro as an alternative to

Where Kokoro shows up in our comparisons, and how it ranked.

Kokoro head-to-head

Straight comparisons against the tools people weigh it against.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.