Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Hugging Face launches the Open TTS Leaderboard to standardize evaluation of open-source text-to-speech models using objective metrics, addressing gaps in fragmented and unstandardized current practices.
The Open TTS Leaderboard introduces objective metrics like ASR-based WER and speaker similarity to evaluate text-to-speech models, reducing evaluation time from weeks to hours. These metrics assess intelligibility and voice identity preservation but do not directly measure naturalness or listener preference. The leaderboard complements, rather than replaces, human preference rankings such as MOS or MUSHRA.
Models are ranked by macro-average WER on English datasets, with hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro leading the pack. Pareto plots visualize trade-offs between WER, batched inference speed, and model size. Performance in English does not necessarily translate to other languages, requiring separate multilingual evaluations.
Multilingual performance is assessed using Seed TTS Eval and CV3 Eval datasets, with character error rate (CER) reported for character-based languages like Chinese, Japanese, and Korean. Strong multilingual models include k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512.
The leaderboard includes a 'Listen' tab for comparing model outputs and a 'Streaming' tab ranking models by time-to-first-audio (TTFA), critical for interactive applications. Streaming performance is measured on H200 GPU and CPU, with kyutai/pocket-tts excelling in both environments.