Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Hugging Face launches Open TTS Leaderboard for multilingual speech models

Hugging Face has launched an automated leaderboard comparing open text-to-speech models across languages, voice cloning, throughput and streaming latency, but warns that the scores do not measure naturalness or listener preference.

D
Sep 30, 2026 · 3 min read

Hugging Face has launched an Open TTS Leaderboard that compares open-source, multilingual text-to-speech and voice-cloning models using automated measures of intelligibility, speed and speaker similarity. It gives developers moving Hugging Face models toward deployment a shared test surface for narrowing the field. The launch authors caution that its scores are proxies, not substitutes for human listening tests.

Hugging Face says its Hub contains more than 8,000 text-to-speech models. The company said automated evaluation can reduce the time needed to assess a model from weeks of collecting human votes to a few hours, allowing more open models to be compared under a consistent setup.

For intelligibility, the system transcribes generated audio with Qwen3-ASR and compares that transcript with the original prompt. It reports word error rate for most languages and character error rate for Chinese, Japanese and Korean, with lower error rates indicating that more of the prompted text was reproduced correctly. The default English view ranks models by macro-average word error rate across the English portions of Seed-TTS Eval and zero-shot CV3-Eval. In the September 30 launch snapshot, Hugging Face named Kokoro-82M, Supertonic 3 and Fish Audio S2 Pro as the leading models on that measure.

Selecting several languages at once makes the table macro-average results across those choices. English and Chinese use Seed-TTS Eval and CV3-Eval, while the other listed languages use CV3-Eval. A voice-cloning toggle limits the table to models that support cloning for the selected languages and adds a speaker-similarity score. That score is the cosine similarity between WavLM speaker embeddings for generated audio and a reference clip. Pareto charts show how similarity and error rates trade off against model size and throughput.

Speed is split into two operating conditions. The main table reports inverse real-time factor for batched offline inference on an NVIDIA H200 GPU. The Streaming tab ranks models by median time to first audio for batch-size-one inference, using 50 English CV3-Eval prompts, consistent hardware and each model’s default voice. The first three runs are discarded as warm-up. H200 results are the default, while CPU measurements are available for a smaller group of models.

The Listen tab puts the audio behind the scores within reach. Users can choose a language and dataset, turn voice cloning on or off, select specific models or a random sample, play generated clips and submit feedback. Hugging Face said it may incorporate community votes after collecting more responses.

The launch post is explicit about what those controls cannot resolve. Word error rate estimates whether the words were reproduced, while speaker similarity estimates whether a cloned voice preserves identity. Neither metric directly measures naturalness, expressiveness or listener preference. A model can therefore score well on the automated table without being the voice a listener would favor.

The launch rankings are a September 30, 2026 snapshot rather than permanent standings. Later results versions may change model coverage, scores and row order, so any reported rank needs to remain tied to its results date.

Hugging Face said at launch that it planned to open-source the evaluation scripts soon. Until those scripts are released, the methodology and generated clips provide visibility into the test setup, but the complete evaluation pipeline is not yet available for independent reproduction.

More news