Google introduces Gemini 3.1 Flash TTS for AI voice
Google is rolling out Gemini 3.1 Flash TTS in preview, a synthetic voice model with precise control over tone, pace, accent and emotion. It works in more than 70 languages, supports multiple speakers and includes SynthID watermarks to detect AI-generated audio.

Google is rolling out Gemini 3.1 Flash TTS, a model that converts text to speech with more control over tone, speed, accent and delivery. It is already available in preview for developers, businesses and Google Workspace users.
The model lets you create more natural voices and direct a performance with instructions written in everyday language. For example, you can ask a character to speak more slowly, whisper a line, change attitude halfway through a sentence or respond with a specific accent.
A voice you can direct
The main new feature is audio tags, commands added to the text to adjust how each segment should sound. They do more than select a voice: they let you define the pace, emotion and delivery of the words.
Google AI Studio adds controls for preparing an entire scene:
- Scene direction: describe the setting and give instructions so characters stay in character across multiple turns.
- Speaker control: assign different audio profiles and adjust each character’s tone, pace or accent.
- Changes within a sentence: tags let you modify the delivery at specific moments without changing the voice’s entire configuration.
- Direct export: once the performance is ready, you can export the parameters as code for the Gemini API and reuse them in other projects.
This could be useful for creating audiobooks with multiple characters, voice assistants, dubbing, video games, narrated courses or content adapted for different countries. The difference is that creators can treat the voice as an artistic direction, not just an automated reading of the text.
Quality, languages and cost
According to Google, Gemini 3.1 Flash TTS achieved an Elo score of 1,211 on the Artificial Analysis leaderboard. The system collects thousands of human preferences in blind tests to compare the quality of different synthetic voices.
Artificial Analysis also placed it in its quadrant of the most attractive options because it combines high-quality voice generation with low cost. The tool also supports native multi-speaker dialogue and works in more than 70 languages.
For you, this could mean applications that sound less uniform and adapt better to the language and context. A single platform could narrate content for multiple markets, change the pace for a child audience or use distinct voices in a conversation between characters.
Where it is available
The rollout began in preview across three areas:
- For developers, through the Gemini API and Google AI Studio.
- For businesses, through Vertex AI.
- For Workspace users, through Google Vids.
Because it is still in testing, performance, controls and terms of use may change before broader availability.
All audio generated by the model includes SynthID, an imperceptible watermark embedded in the sound. Its purpose is to enable the detection of AI-generated content and help reduce the risk of synthetic speech being presented as a real recording.
The next thing to watch is not just whether these voices sound natural, but whether the tools can maintain a consistent identity across long conversations, different languages and emotional shifts. That will determine whether Gemini 3.1 Flash TTS is useful for real productions or only for attention-grabbing demos.