AI News
AI News AgentModel releaseGoogle3 min read

Google improves Gemini, its voice and translation AI

Google has updated Gemini 2.5 Flash Native Audio to improve its voice agents, with better instruction following, real-time data retrieval and more coherent conversations. It has also launched a beta for simultaneous translation in Google Translate on Android, supporting more than 70 languages and 2,000 language pairs.

Google has updated Gemini 2.5 Flash Native Audio so its voice agents can better understand instructions, retrieve information in real time and hold more natural conversations. At the same time, it has started testing a simultaneous translation feature that sends translated conversations directly to your earbuds.

Voice agents that follow the thread better

Gemini 2.5 Flash Native Audio is Google's model designed for real-time voice conversations. The new version improves three important areas for people using assistants, customer service systems or workplace tools over the phone.

  • More accurate function calls: the model is better at identifying when it needs to retrieve external data, such as an order or reservation status, and returns that information without breaking the flow of the conversation. In the ComplexFuncBench Audio test, it scored 71.5% when handling tasks with multiple calls and conditions.
  • Better instruction following: its compliance with developer instructions rises from 84% to 90%, according to Google. This helps the agent better follow rules about its response tone, the steps it should take or the data it can provide.
  • More coherent conversations: it more accurately retrieves what was said in earlier turns, which matters when a call lasts several minutes and you do not want to repeat the same information.

In practice, this makes it possible to build agents that manage complete processes, not just answer isolated questions. For example, a financial services assistant could collect information, check the status of an application and explain the next step during the same call.

Where it is available

The model is now generally available in Vertex AI, Google's Cloud platform for developing AI applications. You can also try it as a preview in the Gemini API and Google AI Studio.

Google has also started adding it to Gemini Live and Search Live. This is the first time Search Live has used this type of native audio, allowing you to ask for help in real time, discuss an idea or search for information more naturally through conversation.

Companies such as Shopify, United Wholesale Mortgage and Newo.ai are already using these capabilities for customer service, mortgage processing and virtual receptionists. Google says United Wholesale Mortgage has generated more than 14,000 loans for its partners through its Gemini-based system, although that figure comes from the company's own report and has not been independently measured.

Simultaneous translation in your earbuds

The other new feature is coming to Google Translate as a beta for live voice translation. The tool listens to a conversation and sends the translation to your earbuds, while attempting to preserve the speaker's intonation, rhythm and tone.

It can work in two ways:

  • Translate multiple languages continuously into a single target language, which is useful for following a conversation or the environment around you.
  • Handle a conversation between two people and automatically change the output language depending on who is speaking.

Google says the feature combines Gemini with its audio capabilities to support more than 70 languages and 2,000 language pairs. It can also automatically detect the language, understand multiple languages in the same session and filter out some ambient noise.

For now, the beta is rolling out in the Google Translate app for Android devices in the United States, Mexico and India. You need to connect a pair of earbuds and tap the “Live translate” option. Support for iOS and more regions will come later, according to Google.

What changes for you

The update does not turn every call into a perfect conversation. But it brings voice AI closer to tasks where it previously failed often: remembering context, following complex rules, retrieving data and responding without artificial pauses.

For users, the most visible change may be real-time translation, especially when traveling or talking with people who speak different languages. For businesses, the important development is the arrival of agents that can complete entire processes by voice.

Google plans to keep refining translation based on feedback from this beta and bring it to more products, including the Gemini API, during 2026. Availability by country, operating system and product will be the key detail to watch before assuming these features are open to everyone.

Google improves Gemini, its voice and translation AI | neversleep.ai