Google launches Gemini 3 Flash, faster and cheaper
Google introduces Gemini 3 Flash, a model that combines advanced reasoning, visual analysis and coding with speeds up to three times faster than Gemini 2.5 Pro, according to its tests. Pricing starts at $0.50 per million input tokens, and the model is already available across several developer and business tools.

Google has launched Gemini 3 Flash, an artificial intelligence model designed to respond quickly, handle complex tasks and cost much less than Gemini 3 Pro. The company says it outperforms Gemini 2.5 Pro on several benchmarks and runs three times faster according to measurements by Artificial Analysis.
The idea is simple: offer advanced reasoning, coding and visual analysis without forcing developers to pay the price or wait for the response times of a larger model.
What Gemini 3 Flash offers
Google presents the model as an option for applications that need intelligence and speed at the same time. In its tests, Gemini 3 Flash achieved:
- 90.4% on GPQA Diamond, a doctoral-level scientific reasoning test.
- 33.7% on Humanity’s Last Exam, without using external tools.
- Higher performance than
Gemini 2.5 Proon several benchmarks, according to Google.
The model also includes the Gemini family’s multimodal capabilities. That means it can work with text, images and video, as well as analyze spatial relationships. Google highlights one especially practical feature: it can run code to enlarge, count or edit elements within an image.
For example, an application could analyze a photograph of inventory, count specific objects and return the result without someone having to review it manually.
How much it costs
In the Gemini API and Vertex AI, the announced prices are $0.50 per million input tokens and $3 per million output tokens. Tokens are the units of text the model processes, such as word fragments or symbols.
Input audio costs $1 per million tokens. Google also includes several options for reducing costs in certain scenarios:
- Temporary context storage can cut the price by up to 90% when information is reused.
- The Batch API offers 50% savings for asynchronous tasks, meaning jobs that do not need an immediate response.
- Paying customers receive production-ready usage limits for nearly real-time applications.
In practice, this could make it viable to use the model to review large volumes of documents, process images or generate code repeatedly without making each query too expensive.
Where it can be used
Gemini 3 Flash is being added to the Gemini API through Google AI Studio, Google Antigravity, Gemini CLI, Android Studio and Vertex AI for businesses.
Google has also shared several early use cases:
- Coding: assistance with creating and modifying code through faster workflows.
- Video games: video analysis and the generation of playable plans and code from a single instruction.
- Deepfake detection: Resemble AI says it analyzes multimodal content four times faster than with
Gemini 2.5 Pro. - Legal documents: Harvey uses it to analyze complex texts while maintaining a focus on speed.
These examples come from customers and partners cited by Google, not from a single independent evaluation covering all use cases.
For developers, the main change is economic and operational: they can test tasks that previously required a more expensive model, with more requests per minute and lower latency. For you, that could mean applications that respond faster, assistants capable of reviewing images or documents, and more accessible coding tools.
Google is pushing Gemini 3 Flash toward a specific goal: advanced intelligence without slow responses or high costs. The important question now is how it performs outside benchmarks, especially in terms of accuracy, stability and real-world tasks at scale.