01 What happened
Hugging Face authors say they have built the Open TTS Leaderboard to evaluate text-to-speech models with objective metrics across complementary performance areas. It measures intelligibility with WER/CER using Qwen3 ASR; speed with batched offline RTFx on an H200 GPU and time-to-first-audio for batch-size-1 streaming on an H200 GPU and CPU; and speaker similarity with cosine similarity between WavLM embeddings. CPU results cover a small but growing set of models.
02 Key details
- The default ranking uses macro-average WER on the English splits of Seed TTS Eval and CV3 Eval. Where Seed TTS Eval has no audio, other languages use CV3 Eval alone; Chinese, Japanese and Korean are reported with CER.
- For streaming, each model runs one audio at a time on the same 50 English CV3-Eval prompts, using its default voice and the same hardware. The first three runs are dropped as warm-up, and the median TTFA is reported.
- The authors say objective evaluation reduces a model check from a couple of weeks of vote collection to a couple of hours. They also stress that the leaderboard does not replace human-preference ranking, while WER and speaker similarity do not measure naturalness or expressiveness.
- In the authors’ own benchmark results, Kokoro-82M, supertonic-3 and fishaudio/s2-pro lead on English WER. Those results come from Hugging Face’s runs and have not been independently verified; the authors say the evaluation scripts will soon be open-sourced.
03 Why it matters
For developers, the leaderboard offers a faster and more reproducible way to compare models than collecting votes for weeks. Its metrics are still proxies: WER and speaker similarity do not measure naturalness, expressiveness or listener preference, and the leaderboard does not replace human rankings.
04 Who it matters to
Voice system engineers, model evaluation specialists and teams selecting TTS models.