AI News
AI News AgentModel releaseGoogle3 min read

Google launches Gemini 3.8 Live for voice agents

Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice models that can converse, process images, and run tasks in the background. The first prioritizes scale and efficiency; the second adds reasoning for complex processes and is coming to products such as Search, Gemini Live, and Workspace.

Google has introduced two new Gemini models designed to make speaking with AI more natural and useful: Gemini 3.8 Live, built for fast, cost-efficient conversations, and Gemini 3.8 Live Extended Thinking, designed for complex tasks that require several steps of reasoning.

The difference matters because these models do more than respond when you finish speaking. They can hold a fluid conversation, process images in near real time, and run tools or API calls in the background while continuing to talk with you.

Two models for different needs

Gemini 3.8 Live is built to operate at scale and at a lower cost. Google is targeting companies and developers that want to create voice agents, such as customer service assistants, application interfaces, or systems that can look up information while speaking with the user.

The model can also detect and switch between 97 supported languages during the same conversation. It understands visual information in near real time as well. For example, you could show it a document or image and ask it to explain it without having to stop the conversation.

Gemini 3.8 Live Extended Thinking is aimed at tasks that require more planning. It can reason and speak at the same time, using cues such as “Let me check that…” while it works on a request. During longer processes, it narrates its progress so you do not feel as though the conversation has stalled.

In practice, this could help solve a problem with several steps, review information from different sources, or complete a task inside an application while keeping you informed.

The results Google is reporting

Google says Gemini 3.8 Live Extended Thinking took first place overall in Artificial Analysis's Speech to Speech Quality Index, with a score of 82.6. It also recorded 68.6% on τ-Voice, a voice-based task completion test, and 35.1% on Sierra's τ-Voice banking benchmark.

On Big Bench Audio, an audio reasoning evaluation, it reached 97.7%. Google says these tests were run through the Live API on the Gemini Enterprise Agent Platform and that the model remains competitively priced compared with other advanced models.

Gemini 3.8 Live, meanwhile, ranked second in Speech Agent Arena, according to Google, while also standing out for its cost efficiency. These figures come from evaluations and user preferences. They do not guarantee the same result in every conversation.

What changes for you

The models will arrive in several Google products and in the tools developers use:

  • Gemini 3.8 Live is beginning to become available to developers through the Gemini API and Google AI Studio.
  • Companies can try it in a private preview of Gemini Enterprise. It will also come to Gemini Enterprise for Customer Experience.
  • Users can find it in Search Live.
  • Gemini 3.8 Live Extended Thinking is also coming to the Gemini API and Google AI Studio.
  • For companies, it will be available in private preview in Gemini Enterprise and soon in Customer Experience and Google Workspace for business.
  • For users, it is being added to Gemini Live and Workspace features for Google AI Pro and Ultra subscribers, as well as Gmail and Keep for Google AI subscribers.

Google is also integrating these models into Search and Workspace experiences so you can solve complex tasks by speaking instead of writing instructions step by step.

An audio safety layer

All audio generated by Google's AI products includes an invisible watermark called SynthID. It is embedded in the sound so it can be detected even when people cannot perceive it, a measure designed to help identify AI-generated content and limit misinformation.

The next step will be seeing how these models perform outside demonstrations and in long conversations with noise, interruptions, and ambiguous requests. Google's direction is clear: voice assistants will stop being simple question-and-answer systems and become agents that can listen, reason, use tools, and keep you informed while they do the work.