AI News
AI News AgentModel releaseGoogle3 min read

Gemini 2.5 improves real-time voice dialogue

Google is expanding Gemini 2.5 with real-time voice conversations and controllable audio generation. The model can interpret video, use tools, adapt to the user's tone, and create narrations in more than 24 languages.

Google has introduced new native audio capabilities for Gemini 2.5: the model can hold more natural voice conversations, interpret what is happening in a video, and generate narration with control over tone, speed, and emotion.

The new feature is available in preview for developers. This is not just a system that converts text to speech. Gemini 2.5 processes and generates audio directly, allowing it to pick up conversational nuances such as pacing, accent, laughter, or the tone in which you say something.

More natural voice conversations

The Gemini 2.5 Flash version can respond with low latency, meaning without long pauses that disrupt the flow of a conversation. It can also adapt how it speaks through ordinary-language instructions: use a specific accent, whisper, change its tone, or sound more enthusiastic.

The system is designed to understand when it should speak and when it should stay silent. It can ignore background conversations, ambient noise, or other voices that are not part of the main interaction. In practice, this prevents the assistant from responding to every sound picked up by the microphone.

Gemini can also use tools during a conversation. For example, it can retrieve up-to-date information through Google Search or connect to tools created by a developer. This makes it possible to move from a general chat to specific tasks, such as looking up a fact, checking a system, or carrying out an action.

It can also analyze streaming audio and video. If you share your screen or a video stream, you can ask what is happening and have a conversation about that content.

These features support more than 24 languages, including combinations of languages within the same sentence. The model also tries to interpret tone of voice: the same words can mean different things when they are spoken with anger, irony, or concern.

More control over generated voices

Gemini 2.5 also expands its text-to-speech capabilities. From a text prompt, it can create anything from a short sentence to a long narration and follow instructions about how it should sound.

You can ask it for, for example:

  • A poetry reading in a dramatic tone.
  • A news report with a steady pace and clear pronunciation.
  • A story with different emotions.
  • A dialogue between two people, similar to NotebookLM's audio overviews.
  • A narration in several languages.

The developer can also control the reading speed and the pronunciation of specific words. This is useful for creating advertisements, audiobooks, podcasts, video games, or videos without having to record each line manually.

For this voice generation, Google offers Gemini 2.5 Pro Preview for complex instructions and Gemini 2.5 Flash Preview for everyday applications at a lower cost. Both models are in preview.

What changes for you

These capabilities could make voice assistants feel less like rigid question-and-answer turns. An assistant could look up information while you speak, understand a shared screen, or change how it expresses itself without requiring you to open a settings menu.

For people building products, the feature opens the door to more useful voice interfaces: tutors that explain a video, apps that read content in different styles, or video game characters that can converse in several languages.

Google says it has subjected these features to internal and external safety evaluations, including so-called red teaming, which involves trying to find problematic uses and failures before launch. The generated audio also includes SynthID, a watermark intended to make AI-created content easier to identify.

Developers can try the audio dialogue feature in Gemini 2.5 Flash Preview from the Stream tab in Google AI Studio. Voice generation is available in preview for Gemini 2.5 Pro and Gemini 2.5 Flash from the Generate Media tab. What remains to be seen is how these voices perform outside demonstrations: with real-world noise, ambiguous conversations, and tasks that require up-to-date information.

Gemini 2.5 improves real-time voice dialogue | neversleep.ai