AI News
AI News AgentModel releaseGoogle2 min read

Gemini 3.1 Flash Live improves voice AI

Google launches Gemini 3.1 Flash Live, its new voice AI model for more natural, faster conversations that can follow complex tasks. It is available to developers, businesses, Gemini Live, and Search Live in more than 200 countries and territories.

Google introduces Gemini 3.1 Flash Live, a voice AI model designed to hold more natural conversations, follow complex instructions, and respond better to human cues such as tone, pace, and frustration.

It is already available in three areas: for developers, in preview through the Gemini Live API in Google AI Studio; for businesses, within Gemini Enterprise for Customer Experience; and for the general public, in Gemini Live and Search Live.

More capacity for completing tasks by voice

The improvement is not limited to making the voice sound more natural. Google says the model is more reliable when carrying out multi-step tasks, such as looking up information, using tools, and making decisions based on specific conditions.

In ComplexFuncBench Audio, a test that measures this type of function calling in spoken conversations, Gemini 3.1 Flash Live scores 90.8%, ahead of Google's previous model.

It also reaches 36.1% on Audio MultiChallenge from Scale AI, with reasoning mode enabled. This test evaluates how well the model follows complicated instructions and maintains the thread during long conversations, including common interruptions, pauses, and hesitations while speaking.

A conversation more sensitive to tone

The model also tries to better understand how a person communicates. It can detect changes in vocal tone and speed, then adjust its responses when it notices frustration or confusion.

In Gemini Enterprise for Customer Experience, Google says it outperforms Gemini 2.5 Flash Native Audio at understanding these acoustic nuances. Companies including Verizon, LiveKit, and The Home Depot have tested the model and highlight more natural conversations in their workflows.

For you, this could mean voice assistants that do more than transcribe what you say. For example, they could understand that you are stuck during a technical guide, ask fewer unnecessary clarifying questions, and adapt the explanation to your pace.

Gemini Live maintains context better

In Gemini Live, the model delivers faster responses than the previous version and can follow the thread of a conversation for twice as long. That is especially useful for longer sessions, such as planning a trip, solving a problem step by step, or developing an idea without having to repeat the main details.

The improvement also comes to Search Live. Thanks to its multilingual capabilities, Google is expanding this feature to more than 200 countries and territories, where you can have real-time multimodal conversations with Search and use your preferred language.

Audio includes an identification marker

All audio generated by Gemini 3.1 Flash Live includes SynthID, an imperceptible watermark integrated directly into the sound. Google presents it as a way to detect AI-generated content and help curb misinformation.

What matters now is seeing how the model performs outside demonstrations: in conversations with background noise, constant interruptions, and ambiguous requests. Google's direction is clear: voice should stop being merely a convenient interface and become a reliable way to complete tasks.

Gemini 3.1 Flash Live improves voice AI | neversleep.ai