Gemini 3.1 Flash Live comes to voice agents
Google launches Gemini 3.1 Flash Live to create voice and vision agents that respond with lower latency and better understand noise, tone and instructions. The model supports more than 90 languages and is now available through the Gemini Live API and Google AI Studio.

Google now lets you build voice and vision agents that respond during a conversation with Gemini 3.1 Flash Live, a new model available through the Gemini Live API in Google AI Studio.
The idea is simple: an application should be able to listen to you, understand what is happening around it and respond without the pauses that often make an AI conversation feel artificial. It can also use external tools to look up information or take actions while you speak.
Fewer problems when conversations happen in the real world
Demonstrations in controlled environments rarely reflect how we speak in everyday life. There is traffic, televisions playing, multiple people talking and incomplete sentences. Google says Gemini 3.1 Flash Live is better at distinguishing the relevant voice from background noise.
That could help with an assistant that guides you while you repair an appliance, an application that handles calls or an agent that analyzes a video stream and answers questions about what it is seeing.
According to Google, the model also improves in three important areas:
- Tool use: it can activate external functions and provide information during a live conversation more reliably.
- Instruction following: it follows complex rules more consistently, even when the conversation changes direction.
- Dialogue fluency: it better interprets elements such as tone, pace and vocal emphasis, while also reducing latency, meaning the time between what you say and the response.
Google says these improvements enable more natural conversations than Gemini 2.5 Flash Native Audio, especially when recognizing acoustic nuances. The announcement does not include specific latency figures or independent test results, so final performance will depend on how each developer integrates the model.
More than 90 languages and connections to applications
Gemini 3.1 Flash Live supports multimodal conversations in more than 90 languages. Multimodal means it can work with different types of information, such as audio, video and text, within the same experience.
The Gemini Live API also includes features designed for real products, not just prototypes:
- Function calling and tool use.
- Session management for long conversations.
- Ephemeral tokens to control API access more securely.
- Integrations for scaling WebRTC connections and distributing them through global networks.
For you, the change may show up in assistants that respond with less delay, understand commands better in noisy places or act within an application without making you type every instruction. But that does not mean every app will behave this way automatically: developers still need to design the workflows, connect the tools and set the agent's limits.
The model is available today through the Gemini API and Google AI Studio. The next area to watch will be its use in phone support, visual assistance and live video applications, where speed and the ability to follow instructions matter more than a brilliant response in an isolated conversation.