September 22, 2026
A pleasant voice does not mean accurate pronunciation: TTS arena leader ranks 11th for pronunciation
Artificial Analysis tested reading difficult passages aloud: Sonic 3.6 ranks first in listener preference and eleventh in accuracy at 74,5%.

Artificial Analysis
@artificialanlys
Introducing Artificial Analysis Pronunciation Robustness — a benchmark of how reliably Text to Speech models pronounce difficult text correctly. Google Gemini 3.1 Flash TTS leads with 88,1%, closely followed by SpaceXAI TTS at 87,6% and ElevenLabs Eleven v3 at 85,6%. Current TTS evaluations, including our TTS Arena, capture overall listener preference (for example, how natural a voice sounds), while Word Error Rate checks whether the right words came out. Pronunciation Robustness adds a view of whether those words are pronounced correctly (for example, “St.” is read as “Saint” and “Street” in the sentence “St. Mary's is on Church St.”). This is critical for voice agents in production: they need to correctly say account details, names, amounts, and more so users can trust them, while doing so with the low latency live conversation demands. How Pronunciation Robustness works Each model voices 454 sentences containing 701 target words or phrases across four categories: 1. Context-dependent (words read differently depending on context, such as wound meaning an injury and wound meaning wrapped) 2. Abbreviation expansion (numbers, dates, units, and notation read naturally, such as 6'2", Chapter XVII, or 1 tsp sugar) 3. Exact sequences (codes, paths, email addresses, and identifiers pronounced exactly, such as .env.local or a.chen@ucsf.edu) 4. Standalone terms (brands, geographic names, and technical terms, such as Arkansas, façade, or genre) We send sentences as written, with no normalization on our side beyond each model's default settings. Selected human listeners judge whether each target was pronounced correctly against a predefined set of acceptable variants; we exclude those who fail the attention check. The score is the share of correct judgments, excluding “could not make out” responses; we publish a model once three or more approved listeners have covered 95% of targets. Key results: ➤ Overall leaders: Gemini 3.1 Flash TTS from @GoogleDeepMind leads with 88,1%, followed by TTS from @SpaceXAI at 87,6%, Eleven v3 from @ElevenLabs at 85,6%, v3 Conversational at 84,8%, and Qwen-Audio-3.0-TTS-Plus from @Alibaba at 81,6%. Gemini 3.1 Flash TTS leads the “context-dependent” and “abbreviation expansion” categories, TTS from SpaceXAI leads “exact sequences,” and Qwen-Audio-3.0-TTS-Plus leads “standalone terms.” ➤ Hardest categories: abbreviation expansion (62,4%) and exact sequences (62,9%) trail context-dependent cases (86,2%) and standalone terms (86,1%). We expect scores for abbreviation expansion and exact sequences to rise noticeably on normalized input text, especially for models with normalization off by default. ➤ Preference ≠ pronunciation reliability: the most liked voices are not always the most accurate. Sonic 3.6 ranks #1 in the Provider Voice Arena with 1276 Elo but only #11 in pronunciation reliability at 74,5%, while Gemini 3.1 Flash TTS ranks #9 in the Arena with 1201 Elo and #1 in pronunciation reliability at 88,1%. Details below ⬇️

· 36K views
The voice listeners like most stumbles over an account number and an address.
On 22.09.2026, Artificial Analysis released its Pronunciation Robustness benchmark: 454 sentences, 701 difficult cases, judged by human listeners. It is the third benchmark line for voice engines, and it diverges from the previous two.
What it measures. The preference arena shows which voice sounds more pleasant, while WER checks whether the right words were spoken. The new benchmark asks whether they were pronounced correctly: in the sentence “St. Mary's is on Church St.”, the first abbreviation is read as “Saint,” the second as “street.”
First place goes to Gemini 3.1 Flash TTS with 88,1% correct pronunciations. TTS from SpaceXAI trails by half a point at 87,6%. Eleven v3 from ElevenLabs scored 85,6%, and Qwen-Audio-3.0-TTS-Plus from Alibaba scored 81,6%.
It breaks on small details. Models read abbreviations, numbers, and units correctly in 62,4% of cases, and exact sequences such as .env.local or an email address in 62,9%. On words where context determines the pronunciation and on names such as Arkansas, the score stays above 86%. A production voice agent gets exactly those payment details, amounts, and addresses.
Sonic 3.6 ranks first in the preference arena and eleventh in pronunciation at 74,5%. Gemini 3.1 Flash TTS is the exact opposite: ninth in listener preference and first in accuracy.
The models received text as written, and the benchmark authors expect both weak categories to improve noticeably on cleaned input.
