Google improves Gemini 2.5 for controlling AI voices
Google updates Gemini 2.5 Flash and Pro to generate voices with more control over tone, pacing and style. The models also improve dialogue with multiple speakers and preserve their identities across 24 languages.

Google has updated its Gemini 2.5 Flash and Gemini 2.5 Pro text-to-speech models so their voices follow instructions on tone, pacing and style more closely. The new versions are available in preview and replace the text-to-speech models released in May.
The difference matters when simply reading a text is not enough. An audiobook needs pauses and emotion; an online course must explain complex concepts without speeding up; and a podcast with several characters has to preserve each voice’s identity throughout the conversation.
More control over tone and style
The models now respond more accurately to prompts such as "cheerful and optimistic," "serious and gloomy" or "nervous and increasingly excited." The voice does not just pronounce the words. It tries to interpret the role you give it.
That can help you create a dramatic narrator, a video game character or a virtual assistant with a specific personality. For developers, it means relying less on manual adjustments to achieve a consistent performance.
Pacing adapts to context
Gemini 2.5 TTS also improves speed control. The model can slow down to emphasize an idea, leave space after an important sentence or speed up an action scene. It also follows explicit pacing instructions more closely.
For example, a prompt can ask it to begin telling a mystery story nervously and gradually speed up until it reaches excitement and relief. The goal is for the change to be audible in the performance, not just visible in the written text.
More consistent voices in dialogue
The new versions are designed for podcasts, simulated interviews and stories with multiple characters. Each participant should retain their vocal identity, and transitions between speakers should sound more natural.
Google says this consistency is also maintained across the 24 supported languages, preserving each character’s tone, register and style throughout the conversation. This is especially useful for translated content and courses aimed at audiences in multiple countries.
Two models depending on your priority
Gemini 2.5 Flash TTS: optimized for low-latency responses, meaning less waiting.Gemini 2.5 Pro TTS: focused on higher output quality.
Both are available through the Gemini API in Google AI Studio. Developers can try them in the Playground, build prototypes and consult the documentation, prompting guide and API Cookbook.
The audio platform Wondercraft already uses Gemini TTS for features such as Convo Mode, which generates conversations between multiple voices, and Director Mode, which lets you control pronunciations, intonation and nonverbal sounds.
For you, the change will be most noticeable in applications that turn text into audio experiences: fewer flat voices, easier-to-follow dialogue and more control without having to edit every sentence by hand. What remains to be seen is how these improvements perform outside demos and in long productions, where consistency is harder to maintain.