September 30, 2026
Compare speech generation on the same texts: Open TTS Leaderboard is available in your browser
Hugging Face has brought speech samples and model ratings together in one Space, including voice cloning.
As of Sep 9, OmniVoice leads the Open TTS Leaderboard in pronunciation accuracy across 9 languages, with and without voice cloning.
Equal conditions. Instead of comparing individual demos, Hugging Face gives models the same texts from two datasets: Seed-TTS eval and CV3-Eval. In the browser-based Hugging Face Audio Space, you can listen to and rate samples to choose a voice for an app or narrated content.
There is no single winner for every task. WER measures errors in spoken words, SIM assesses similarity to the original voice, and TTFA shows the time until audio starts.
| Model | Result as of Sep 9, 2026 | | --- | --- | | OmniVoice | 1st in average WER across 9 languages with and without cloning. 2nd in average SIM. 1st in SIM for English, French and Korean | | Fish Audio S2 Pro | 2nd in average WER across 9 languages. 1st in WER for German | | Qwen3-TTS-GGUF | 1st in time to first audio in a single-request streaming test |
Running OmniVoice. Running it locally requires PyTorch and torchaudio 2.8.0, followed by installation with `pip install omnivoice`. Load the model with `OmniVoice.from_pretrained("k2-fsa/OmniVoice")` and run voice cloning with `generate(text=..., ref_audio=..., ref_text=...)`, passing the new text, a voice sample and the sample's transcript.
Original source: [Hugging Face, Sep 30, 2026](https://huggingface.co/blog/open-tts-leaderboard).
OmniVoice's code is available under Apache 2.0, but its weights use a CC-BY-NC license that does not allow commercial use.
