Hugging Face has launched the Open TTS Leaderboard, a new evaluation framework designed to keep pace with the rapid release of open-source text-to-speech (TTS) models. As of September 30, 2026, over 8,000 TTS models are available on the Hugging Face Hub, yet evaluation methods have remained fragmented.
Product LaunchesHugging FaceOpen TTS Leaderboard
Hugging Face Launches Open TTS Leaderboard for Scalable Multilingual Model Evaluation
While human preference scores like MOS (Mean Opinion Score) are considered the gold standard, arena-style leaderboards often struggle to scale or lack representation of open-weights models due to hosting requirements. The Open TTS Leaderboard addresses these limitations by using objective metrics that can be computed in hours rather than weeks.
The leaderboard utilizes Word Error Rate (WER) as a proxy for intelligibility and speaker similarity estimates to measure voice identity preservation. For character-based languages such as Chinese, Japanese, and Korean, Character Error Rate (CER) is reported. The platform also features a "Voice Cloning" mode and a "Streaming" tab that ranks models based on Time-to-First-Audio (TTFA), a crucial metric for interactive voice agents.
Currently, English performance is ranked using the Seed TTS Eval and CV3 Eval benchmarks, while multilingual rankings allow users to toggle between different languages. Hugging Face intends for the leaderboard to be community-driven and plans to open-source the evaluation scripts on GitHub.
Sources
- Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning (Hugging Face Blog, 2026-09-30)