SlopTV

news

Google's Gemini 3.8 Flash TTS claims #1 on Hume's Voice Design Benchmark at 71.4, with ElevenLabs still ahead on two subscores

The two new speech models ship into the Gemini API, Notebook and Google Vids, with placements cited on Hume AI's benchmark and Voice Arena but no vote counts attached.

By Daniel Ochoa · · Updated

Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. Google reports Flash TTS at 71.4 overall on Hume AI's Voice Design Benchmark, first place, with 60.8 on accent modeling. ElevenLabs scored higher on Hume's individual voice quality and human-like variation categories.

Google released two text-to-speech models on September 23, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, and attached a leaderboard number to both. Flash TTS is reported at 71.4 overall on Hume AI's Voice Design Benchmark, which Google presents as the top position, alongside an accent modeling score of 60.8. The detail that does not travel with the headline is that ElevenLabs still scored higher in Hume's individual voice quality and human-like variation categories, so the claim rests on a composite index rather than a clean sweep of the categories inside it.

Prompt a voice into existence, then direct it line by line

The functional change is that a voice is now something you describe rather than something you pick from a list. Google says Flash TTS can build an original voice from a natural-language description, supports more than 100 languages and dialects, and exposes over 2,000 production voices, up from the 30 originals in the previous generation. Voice replication recreates a consistent voice from a 30-second sample, and Google requires a matching consent recording to go with it. That feature is not available in Google AI Studio in several regions, including the UK, the EEA and India.

Both models take stage directions at the level of individual lines, covering pacing, dialect shifts, emotional tone and backchanneling, and both support two-speaker scene staging from a single script. Google also claims timbre and pacing hold across hours of continuous audio with minimal speaker drift, which is the kind of claim that only shows up as a problem at production length.

What the benchmark placement covers, and what it leaves out

Hume AI's Voice Design Benchmark is a third-party scoreboard, but Hume is itself a voice AI company selling into the same market, which is worth holding in mind when a rival's model is announced as first on it. Google reports Flash TTS first overall and Flash-Lite second on Hume's overall quality index, and says both models took top or near-top positions in blind human preference tests on Voice Arena in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.

Those are placements, not sample sizes. The announcement names the languages and the rankings without saying how many blind comparisons produced them, which is the same gap that makes public arena rankings for generated media hard to read, as with the video leaderboard ranking 12 models on 1,394 total votes. A first-place finish on a few hundred pairwise votes and a first-place finish on a few hundred thousand are different facts wearing the same label.

Where the models land, including Google Vids

Flash TTS is rolling out in the Gemini API and Google AI Studio for developers and in Gemini Notebook for general users, with Gemini Enterprise access listed as coming soon. Flash-Lite TTS is positioned for high-volume, cost-sensitive work such as dubbing and voice agents, and reaches consumers through Google Vids, the same product Google recently stopped paywalling its AI video generation in.

Google named Agora, LiveKit, Pipecat and Vercel as developer platforms building on the speech generation path, and Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as companies integrating the models for dubbing and media localization. That partner list is the real signal about placement. Video models such as Veo 3.1 already generate their own synchronized audio in a single pass, so a standalone TTS stack is aimed less at replacing that and more at the localization and revoicing work that happens after a clip exists.

Sources: Google, Android Authority, MarkTechPost, FoneArena.


DO

Daniel Ochoa: Covers model launches, shutdowns and pricing changes as they happen. Reads deprecation notices for a living so you do not have to.