Google launches Gemini 3 Flash, its fastest AI
Google launches Gemini 3 Flash, a model that combines Gemini 3 Pro's reasoning with lower latency and cost. It replaces Gemini 2.5 Flash in the Gemini app and is starting to arrive in AI Mode, while developers can use it through the API and other Google tools.

Google has launched Gemini 3 Flash, an artificial intelligence model designed to offer capabilities close to those of larger models, with faster responses and lower costs. It is now starting to roll out to the Gemini app, Search and developer tools.
The company presents it as a more efficient version of Gemini 3 Pro. It retains the ability to reason, understand images, audio and video, and carry out multi-step tasks, but it is designed to respond with less delay and use fewer resources.
What changes for you
Gemini 3 Flash replaces Gemini 2.5 Flash as the default model in the Gemini app. Access is rolling out globally at no cost, although availability may vary by country and product.
In practice, you can use it for tasks such as:
- Analyzing a video and turning it into an action plan.
- Understanding an image, document or audio recording.
- Creating a simple app by describing it with your voice, without knowing how to code.
- Solving complex questions with organized answers and useful links.
- Planning a trip around a budget, schedule and specific places.
Google is also starting to make it the default model in AI Mode, its AI-powered search mode. There, it can combine the model's reasoning with up-to-date web information, local data and recommendations.
The goal is for Search to do more than answer a question. It should also help you complete an objective. For example, it can organize research on a difficult topic or prepare an itinerary with multiple conditions.
Faster and cheaper for building products
For developers, the main difference is speed and price. Google says Gemini 3 Flash outperforms Gemini 2.5 Pro and runs three times faster, according to measurements by Artificial Analysis.
Its announced API price is $0.50 per million input tokens and $3 per million output tokens. Tokens are the small units of text the model processes. Input audio costs $1 per million tokens.
The company also says it uses 30% fewer tokens on average than Gemini 2.5 Pro in typical traffic, while delivering stronger performance on everyday tasks. For difficult queries, it can spend more time reasoning. For simple ones, it responds without using unnecessary resources.
That matters when an application needs to respond many times per second. An assistant inside a video game, a tool that analyzes videos or a system that reviews documents can feel smoother and cost less to operate.
What performance it offers
According to tests shared by Google, Gemini 3 Flash achieves these results:
- 90.4% on GPQA Diamond, an evaluation of advanced scientific questions.
- 33.7% on Humanity’s Last Exam, without using external tools.
- 81.2% on MMMU Pro, which measures understanding of multimodal content, such as images and text.
- 78% on SWE-bench Verified, a programming test based on real software tasks.
These results were published by Google and should be treated as benchmark measurements, not a guarantee for every use case. Even so, the most notable figure is that Gemini 3 Flash outperforms Gemini 3 Pro on SWE-bench Verified, according to the company.
Gemini 3 Flash is already available in preview for the Google AI Studio API, Google Antigravity, Vertex AI, Gemini Enterprise, Gemini CLI and Android Studio. Companies including JetBrains, Bridgewater Associates and Figma are also among its first users.
The direction set by this launch is clear: AI models are no longer competing only to answer more difficult questions. They are also competing to do it quickly, at a cost that makes it possible to use them millions of times a day and inside products that need to react instantly. What matters now is how Gemini 3 Flash performs outside benchmarks and how much of this rollout actually reaches each user.