AI News
AI News AgentModel releaseGoogle4 min read

Google introduces Gemini 2.0 for AI agents

Google launches Gemini 2.0 Flash experimental, a model that can use tools, generate text, images and audio, and plan tasks under supervision. It also introduces Deep Research and agent prototypes for browsing websites, programming and assisting in the real world.

Google introduces Gemini 2.0, a new family of models designed so AI can do more than answer questions. It can also plan tasks, use tools and act under your supervision.

The first release is Gemini 2.0 Flash experimental, a fast model that Google has offered to developers and Gemini users since December 11, 2024. The company says it outperforms Gemini 1.5 Pro on key benchmarks and responds at twice the speed, although feature availability varies by access type.

From answering questions to getting things done

The central idea behind Gemini 2.0 is AI agents. These are systems that can understand a situation, break a task into steps and carry out some actions on your behalf, while still asking for your permission when necessary.

For example, an agent could look up information, write code or interact with a webpage to complete a process. This does not mean giving it full control of your computer: Google says its prototypes keep the user in the loop and request confirmation before sensitive actions, such as making a purchase.

Gemini 2.0 Flash also includes multimodal capabilities. That means it can receive text, images, video and audio, and that the model is designed to generate different types of responses:

  • Text combined with generated images.
  • Multilingual audio through text-to-speech conversion.
  • Direct use of Google Search, code execution and functions created by developers.
  • Real-time interaction with audio and video through a new API.

Image generation, audio and some advanced features are initially available to selected partners only. Google expects to expand availability in January 2025, along with other model sizes.

What you can try now

The experimental version of Gemini 2.0 Flash is available to developers through the Gemini API in Google AI Studio and Vertex AI. In the Gemini app, users worldwide can select it from the model menu on the web for desktop and mobile, while its arrival in the mobile app is described as imminent.

Google is also enabling Deep Research for Gemini Advanced subscribers. This feature acts as a research assistant: it explores complex topics and prepares reports using advanced reasoning and an extensive context window, meaning it can work with large amounts of information within a single task.

In Search, Google is testing Gemini 2.0's reasoning capabilities in AI Overviews, its AI-generated summaries. The company says the feature already reaches 1 billion people and that the new version will be able to handle multi-step questions, advanced mathematics, image-based queries and programming. A broader rollout is planned for early 2025.

The agents Google is testing

Google DeepMind is also introducing several prototypes that are not yet general-purpose products:

  • Project Astra explores a universal assistant that can converse in multiple languages, use Search, Lens and Maps, and remember up to 10 minutes of a session. It also aims to respond with latency similar to a human conversation.
  • Project Mariner works as an experimental Chrome extension that interprets what appears on a webpage and can click, type or scroll to complete tasks. It scored 83.5% on the WebVoyager benchmark, according to Google, although the company acknowledges that it can still be slow and inaccurate.
  • Jules is a coding agent that integrates with GitHub, analyzes a problem, proposes a plan and executes changes under the developer's direction.

Google is also researching agents for video games and robotics. In both cases, the technology remains at an early stage and is presented as an experiment, not as a feature ready for everyday use.

Trust is still the limit

The more a system can do for you, the more important it becomes to control its errors. Google says it is testing these agents with selected users, external experts and safety evaluations. It is also working to help them detect malicious instructions hidden in webpages, emails or documents, a risk known as prompt injection.

For you, the immediate change is limited: you can try a faster model and research features, but agents capable of browsing websites or carrying out tasks are still being tested. The important point about Gemini 2.0 is the direction it sets: AI is no longer just a tool for conversation and is starting to become an interface that can take action. The question to watch is whether it can learn to do so with enough accuracy, transparency and human control.