AI News
AI News AgentModel releaseGoogle3 min read

Gemini 3.8 debuts voices with creative control

Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS, two models for creating custom voices and directing their performance line by line. They support more than 100 languages, consent-based voice replication, and tools for audiobooks, dubbing, podcasts, and voice agents.

Google introduces two text-to-speech models in the Gemini 3.8 family that can create voices from scratch, replicate a voice with authorization, and direct every line with precise instructions. The goal is to move from choosing a preset voice to controlling an entire performance.

Gemini 3.8 Flash TTS is designed for creating characters and performances. You can describe a voice's accent, role, and characteristics in natural language, then specify line by line how it should speak: with pauses, changes in pace, whispers, dialects, or reactions such as laughter and sighs.

Gemini 3.8 Flash-Lite TTS has a different goal: generating large volumes of audio more cost-efficiently. Google is targeting tasks such as dubbing, content creation, and voice agents that need to respond expressively at scale.

From a preset voice to a vocal studio

The Flash TTS model expands the available library from 30 original voices to more than 2,000 production-ready voices. It includes regional varieties such as Mexican Spanish, Quebec French, and Scottish English, along with support for more than 100 languages and dialects.

It also lets you create a custom voice by describing it with text. For example, you can request the voice of a narrator with a specific regional accent or design a fictional character with a particular way of speaking.

Voice replication works from a 30-second voice sample, but only when the user is authorized to use it. Google says the system requires a spoken consent recording from the voice owner and adds SynthID watermarks and C2PA credentials to help identify generated content.

The option to blend and adjust voices from the library by modifying characteristics such as timbre, pitch, speed, or accent will arrive later. Custom voices can also be saved to maintain a consistent performance across episodes, videos, or projects.

Line-by-line control and two-voice conversations

Both models let you add performance instructions directly to the script. The same scene can call for a calm delivery in one line and a suspenseful tone in the next, without having to generate the entire audio track with a single configuration.

Their features include:

  • Long-form audio generation with stable pacing and timbre for hours.
  • Scenes with two speakers from a single script.
  • Distinct conversation turns between the two voices.
  • Nonverbal cues such as <laughs>, <sigh>, or <gasp>.
  • Active-listening interjections such as <mhm> or <yeah> to make dialogue sound more natural.

In practice, this could help produce an audiobook with consistent characters, create a two-voice podcast, or dub a video while preserving regional nuances. It could also give voice agents handling real-time conversations more control.

The results Google reports

Google says Gemini 3.8 Flash TTS ranks first overall on Hume AI's Voice Design Benchmark, with a score of 71.4, and also leads the accent modeling evaluation, with 60.8.

On Hume AI's overall quality index, the Flash TTS and Flash-Lite TTS models rank first and second, respectively, according to the company. Google also says both models achieved strong positions in blind Voice Arena evaluations in languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

These are results reported by Google from specific tests, not a guarantee that every voice or language will sound equally good in every situation.

Where you can try them

The models are starting to roll out in Google AI Studio and the Gemini API for developers. The Flash TTS model is also being added to Gemini Notebook for general users, while Flash-Lite TTS is coming to Google Vids.

For businesses, both models will soon be available through the Gemini Enterprise API. Google is also working with platforms such as Agora, LiveKit, Pipecat, and Vercel, as well as dubbing and content creation companies including HeyGen, Wondercraft, and Ollang.

The important change is not just that Gemini can read text aloud. It is that you can treat voice as an editable part of production: you can design who is speaking, direct how they interpret every line, and maintain that style throughout a project. The question now is whether its consent and detection controls will be sufficient as these voices begin to be used in commercial content and automated conversations at scale.

Gemini 3.8 debuts voices with creative control | neversleep.ai