September 23, 2026
Google voice generation joins the leaders: Gemini 3.8 Flash TTS ranks No. 2 in Arena
On September 23, Gemini 3.8 Flash TTS took #2 in Provider Voice Arena with 1260 Elo and costs $33,0 per 1 million characters.

Artificial Analysis
@artificialanlys
Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Gemini 3.8 Flash TTS immediately took first place in our Pronunciation Robustness benchmark and second in the Provider Voice Arena ranking. Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google DeepMind's new speech synthesis models. They support ready-made voices and voice cloning. Gemini 3.1 Flash TTS led our Pronunciation Robustness at launch. Now Gemini 3.8 Flash TTS has taken first place and surpassed it by 64 Elo in Provider Voice Arena. Key points: ➤ #2 in Provider Voice Arena: Gemini 3.8 Flash TTS debuted in second place with 1263 Elo. It trails Sonic 3.6 with 1272 Elo and leads Qwen-Audio-3.0-TTS-Plus with 1260 Elo. Gemini 3.8 Flash-Lite TTS debuted in sixth place with 1236 Elo, 1 point behind Simba 3.2. Both models rank above Gemini 3.1 Flash TTS with 1199 Elo. ➤ #1 in Pronunciation Robustness: Gemini 3.8 Flash TTS took first place with 89,5%. It is followed by Gemini 3.1 Flash TTS with 88,2%, SpaceXAI TTS with 87,6%, and Gemini 3.8 Flash-Lite TTS with 87,4%. Three Google models made the top four. ➤ Multilingual performance: across nine language-specific Controlled Voice Arenas, Gemini 3.8 Flash-Lite TTS debuted first in Japanese, second in Arabic, and third in German. Gemini 3.8 Flash TTS took second place in Japanese and Portuguese. ➤ Speed: Gemini 3.8 Flash TTS processes 44,1 characters per second, Gemini 3.8 Flash-Lite TTS — 40,2. That is about 2,7 and 2,4 times faster than real-time speech. ➤ Price: Gemini 3.8 Flash TTS costs $32,98 per 1 million characters, Gemini 3.8 Flash-Lite TTS — $22,07. Both models cost more than Gemini 3.1 Flash TTS at $18,31, but are notably cheaper than Eleven v3 at $100. More details below ⬇️
· 25.6K views
Gemini 3.1 Flash TTS previously led Pronunciation Robustness with 88,2%. Now Gemini 3.8 Flash TTS has taken first place with 89,5%, while Google recommends Flash-Lite instead of Gemini 3.1 Flash TTS.
Choosing between versions. Flash processes 44,1 characters per second versus Lite's 40,2. Lite costs $22,1 per 1 million characters, Flash $33,0. Both models support ready-made voices and voice cloning. Flash supports 130 languages, Lite 101.
For single-voice speech generation, the Interactions API receives the `gemini-3.8-flash-tts` model, verbatim text in `input`, style in `speech_metadata`, and a voice in `generation_config.speech_config`. By default, the API returns WAV. Flash has a limit of 8192 input and 16 384 output tokens, with Batch API, Flex, and Priority inference available.
Voice design requires Google GenAI SDK `google-genai` 2.25.0+ or `@google/genai` 2.24.0+, or `POST /v1beta/voices`. Voice cloning requires a clean 10–30 second recording and a separate recording of the person's consent. Stateful mode stores up to 200 voices per project per year, while a stateless key is valid for 7 days.
In Pronunciation Robustness, people check whether the model pronounced English fragments correctly. All models read the same texts in a fixed voice, and results are published after 95% of clips have been reviewed by at least three independent people.
Google already recommends Flash-Lite as a replacement for Gemini 3.1 Flash TTS.

