Google introduces Gemini 3.1 Flash-Lite for scale
Google introduces Gemini 3.1 Flash-Lite, a preview model designed for workloads with many requests. It promises faster responses at a lower cost, with prices starting at $0.25 per million input tokens.

Google has introduced Gemini 3.1 Flash-Lite, a model designed to process large volumes of tasks at a lower cost and with faster response times. Starting today, it is available in preview for developers through the Gemini API, Google AI Studio and Vertex AI for businesses.
The model is built for products that need to respond many times a day, such as automatic translation, content moderation and real-time assistants. It costs $0.25 per million input tokens and $1.50 per million generated tokens. Tokens are the small units of text the model processes.
More speed at a lower cost
According to Google, Gemini 3.1 Flash-Lite outperforms Gemini 2.5 Flash in speed while maintaining similar or higher quality. In the Artificial Analysis benchmark, it takes 2.5 times less time to deliver the first response token and generates text 45% faster.
That first result matters in interactive applications. For example, if a tool reviews hundreds of messages or generates responses for users, lower latency can shorten the wait and make the total cost easier to control.
Google also cites a score of 1432 Elo points in the Arena.ai ranking. In other tests, the model scored 86.9% on GPQA Diamond, which measures advanced reasoning, and 76.8% on MMMU Pro, focused on understanding text and images.
The developer decides how much the model reasons
Gemini 3.1 Flash-Lite includes configurable reasoning levels in AI Studio and Vertex AI. In practice, developers can choose how much effort the model devotes to a task: less for fast, affordable responses, or more for problems that need detailed analysis.
That option makes it possible to use the same model for very different tasks:
- Translate large volumes of content.
- Classify and moderate posts.
- Create interfaces and dashboards from instructions.
- Generate simulations or solve multi-step tasks.
The difference matters because not every request needs the same level of analysis. Routine moderation can prioritize speed, while generating a functional dashboard may require more reasoning.
What this means for you
If you use an application that integrates Gemini, this launch does not mean the model will automatically appear in every service. For now, it is available in preview for developers and businesses, which will have to decide whether it fits their products and accept the conditions of this phase.
For people building AI tools, the proposition is straightforward: run more requests with a model from the Gemini 3 family without paying the price of a larger option. Google says early testers, including Latitude, Cartwheel and Whering, are already using it to handle complex tasks at scale.
The key question is whether this combination of price, speed and quality holds up when the model leaves preview and faces real workloads. The move confirms a clear trend: competition is no longer focused only on creating more capable models, but also on making them usable thousands or millions of times without costs and wait times spiraling.