AI News
AI News AgentModel releaseGoogle5 min read

Google introduces Gemini 4 Argon for complex tasks

Google introduces Gemini 4 Argon, a model capable of maintaining complex tasks with up to 1 million tokens of context. It will first reach cybersecurity defenders and selected customers while the company strengthens its safety systems before expanding access.

Google introduces Gemini 4 Argon, an AI model designed to maintain its reasoning across long, complex tasks. For now, it is not available to the public: it is being tested with selected cybersecurity defenders through the Fairwind program.

The company plans to expand access gradually, starting with paying API customers and Google AI Ultra subscribers. Before that, it wants to strengthen its safety systems and collect results from the first evaluators.

A model built for multistep work

Argon can work with up to 1 million tokens, the unit used to measure the text a model processes and generates. That is a significant increase over the previous limit of 64,000 tokens, allowing it to analyze large amounts of information or work on a task for much longer without losing context.

In practice, that points to tasks such as migrating hundreds of thousands of lines of code, investigating a legal issue using many documents or analyzing an entire computer system for flaws. The promise is not just better answers, but completing processes from beginning to end with less human intervention.

Google says the initial price will be $2 per million input tokens and $10 per million output tokens. Cached input tokens will receive a 95% discount from the standard rate.

What it is doing inside Google

Thousands of Google employees already use Argon for programming, research and writing tasks. The company highlights several internal results, although some still depend on later deployments and reviews:

  • In quantum algorithm optimization, it outperformed a published benchmark by 40% by reducing the resources needed for certain calculations.
  • Its agents analyzed memory usage data in Google's data centers and identified optimizations that could free up more than 300 tebibytes of memory once implemented. The total estimate ranges from 500 tebibytes to 1 pebibyte.
  • It is helping migrate C and C++ code to Rust, a language designed to reduce certain memory errors. The projects range from libraries with tens of thousands of lines to the Zircon kernel of the Fuchsia operating system, with more than 800,000 lines.

In the libgav1 video decoder, the agents replaced 32,000 lines of specialized code with memory-safe Rust. Google says the result is 2.7 times faster than the previous Rust version, produces exactly the same video and comes close to the performance of the original C++ code.

These changes do not go directly into production. The company says it subjects them to automated and manual audits, emulation tests and additional reviews because they involve critical systems.

Results in programming and business work

Google says Gemini 4 Argon scores 77.9% on DeepSWE v1.1, an evaluation of real-world software engineering tasks that require multiple steps. It also leads the Vals index, which measures the potential economic impact of models in finance, programming, law and taxation.

On AutomationBench, a Zapier test that evaluates whether an AI can execute complete business processes, Argon scores 51.3% and ranks first according to Google. In long-video understanding, it achieves 91.7% on LVBench, a test that measures whether a model can find relevant information in lengthy recordings.

The idea is that it can combine text, code, images, charts, videos and documents to perform professional tasks. For example, it could analyze a financial report and its charts, review a long video and then act on the findings.

A special focus on cybersecurity

Argon has also been trained to locate, verify and fix software vulnerabilities. Google is releasing it without the usual cyber restrictions for its own teams and trusted defenders, who will be able to use its full capabilities for protection tasks.

Wiz, a cybersecurity company, already uses it through Scan for Good, a program that searches for and fixes serious risks in public infrastructure at no charge. In an initial test, Argon identified a critical vulnerability exposing personal data in healthcare software used by hospitals in several countries. According to Wiz, other advanced models had not detected the risk.

On CWE-bench v1, a vulnerability-fixing evaluation, it ties for first place with a score of 68%. Google also says it outperforms Gemini 3.8 Flash Cyber in penetration tests without access to source code, where it must analyze live web systems, discover their attack surface and demonstrate that the identified flaws are real.

Why the launch will be gradual

A model capable of investigating code, finding vulnerabilities and acting on systems can also be used to cause harm. That is why Google is testing several safeguards before offering it more broadly:

  • Blocking requests related to cyberattacks and chemical, biological, radiological or nuclear risks without preventing legitimate scientific research.
  • Protection against indirect prompt injection, an attack that inserts malicious instructions into documents or pages to divert the model's behavior.
  • Monitoring the system's reasoning and actions so it can be stopped if it tries to go beyond what the user requested.
  • Isolated, hardened environments for training and evaluating agents capable of taking actions.

Google also says its internal and external teams have tried to break these safeguards through manual and automated attacks. The company is also participating in the voluntary U.S. government process to provide early access to advanced models before their release.

For you, the immediate impact is limited: Argon is not yet an open chatbot for everyone. What matters is the direction it points to. Google is preparing models that do more than answer questions: they sustain long projects, handle large volumes of information and make changes to real systems. The next step will be finding out whether that autonomy maintains its accuracy and safety when access is no longer restricted to evaluators and specialized customers.