Summary

Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS for custom voice creation and directed speech generation. The company says the models support more than 100 languages and dialects, with voice replication safeguards including consent verification and SynthID watermarking.

Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026, adding two text-to-speech models focused on customizable voices and more directed performances. The company says developers can use them through Google AI Studio and the Gemini API, with experiences also planned across Gemini Enterprise, Gemini Notebook and Google Vids. Google AI Studio is available for developers to try the capabilities starting today, according to the announcement.

Custom voices and directed dialogue

Gemini 3.8 Flash TTS is designed for creating voices from natural-language descriptions, with controls for characteristics such as role and accent. Google says it supports more than 100 languages and dialects and provides access to a library of more than 2,000 production-ready voices. The model can also replicate a voice from a 30-second audio sample, provided the user has the right to use it.

Both models allow creators to guide delivery line by line using script cues for elements such as pacing, emotion and dialect. Google describes support for two-speaker scenes, conversational interjections and vocal sounds such as laughter or sighs. It says Flash TTS can maintain voice quality and character across hours of generated audio with minimal speaker drift. Flash-Lite is positioned for high-volume uses such as dubbing, audio production and voice agents, with controls for tone, pacing and expressive nuance.

Google says voice replication requires a verbal consent recording from the voice owner that matches the reference speaker. The company also says every audio clip generated by its Gemini Audio models carries an imperceptible SynthID watermark intended to make AI-generated speech detectable. C2PA credentials are listed among the protections associated with voice replication. Saving designed voices is intended to help keep them consistent across projects. Voice remixing—adjusting qualities such as pitch, timbre, pace or accent in an existing library voice—is described as coming soon.

Company-reported evaluation results

Google reports that Gemini 3.8 Flash TTS ranked first on Hume AI’s Voice Design Benchmark, with a score of 71.4, and in accent modelling, with a score of 60.8. It also says Flash TTS and Flash-Lite TTS ranked first and second, respectively, on Hume AI’s Overall Quality Index.

The company further reports top positions for both models in blind human-preference evaluations on Voice Arena across languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. These results are presented by Google as evidence of voice-design, audio-quality and multilingual performance; the announcement does not describe the evaluation procedures in detail.

Sources