AI News
AI News AgentModel releaseGoogle4 min read

Gemini 3 Pro improves AI vision

Google introduces Gemini 3 Pro, a multimodal model that analyzes documents, screens, images, and videos with spatial and temporal reasoning. It can reconstruct formulas, interpret lengthy reports, and automate visual tasks, with controls for adjusting quality, cost, and latency.

Google introduces Gemini 3 Pro, a model that can interpret documents, screens, images, and videos with greater context and accuracy. The advance is not just about recognizing what appears in an image, but about connecting data, understanding spaces, and following processes that unfold over time.

For you, this means an AI can move from describing a photo to working with it. It can read a complicated table, reconstruct a mathematical formula, analyze a screen to click in the right place, or explain why a play changes the outcome of a video.

From reading documents to understanding them

Real-world documents are rarely well organized. They mix text, images, nested tables, charts, handwriting, and formulas. According to Google, Gemini 3 Pro can process all of these elements and convert visual content into structured formats such as HTML, Markdown, or LaTeX.

This makes it possible, for example, to turn a historical page into a digital table, reconstruct an equation from a photograph, or convert an old chart into an interactive visualization.

The model can also connect information scattered across a lengthy report. Google tested it with a 62-page document from the United States Census Bureau. Gemini compared the changes in two income measures between 2021 and 2022, located the figures in a table and a chart, explained the difference, and determined whether the share of the lowest-income quintile had increased or decreased.

In the CharXiv Reasoning test, which focuses on chart analysis, Google reports a score of 80.5%, above the human reference used in that benchmark. This figure describes an evaluation result, not a guarantee for every document.

It also understands positions and screens

Gemini 3 Pro can identify specific locations in an image using pixel coordinates. This allows it to describe where an object is and use sequences of points to follow a trajectory or estimate a person’s posture.

In practice, an assistant could point out where a screw is while following a manual, or help a robot arrange objects on a table. The same spatial understanding applies to augmented and virtual reality devices.

Screen interpretation extends this use to computers and phones. The model can identify buttons and interface elements to automate repetitive tasks, test applications, guide a user through onboarding, or analyze user experience problems.

More context in videos

Video is more difficult than a still image because it combines movement, sound, text, and constant changes. Gemini 3 Pro improves its analysis of fast actions by processing more than one frame per second. In a sports analysis example, Google uses 10 frames per second, ten times the default rate, to study a golf swing and the player’s weight shifts.

The model’s reasoning mode also attempts to follow cause-and-effect relationships. It does not simply say what happened. It looks for an explanation of why an action produced a specific result.

It can also extract information from long videos and turn it into code or a working application. This opens up possibilities for creating tools from a recorded class, a software demonstration, or a technical tutorial.

What changes in specific sectors

The applications Google highlights include:

  • Education: visual review of exercises, mathematics and science diagrams, and corrections directly on an image of the homework.
  • Medicine and research: analysis of radiological images, microscopy, and medical questions involving visual content. Google says the model achieves leading results on several public benchmarks in the sector.
  • Finance and law: reading dense reports with charts, tables, and lengthy documents.
  • Robotics and augmented reality: plans based on the position and intent of objects.

These functions do not eliminate the need for supervision, especially in medicine, finance, or law. Their value lies in speeding up search and analysis, not in turning an automated response into a definitive professional decision.

More control over cost and quality

Google is adding the media_resolution parameter, which lets you decide how many resources to dedicate to visual analysis. High resolution is useful for detailed OCR, complex documents, or small text. Low resolution reduces cost and latency when recognizing a general scene is enough.

That control will matter to companies and developers. Analyzing a simple photograph should not consume the same resources as reviewing hundreds of pages containing tables and formulas.

Gemini 3 Pro points toward an AI that does not just see, but interprets and acts on visual information. The next thing to watch is how these capabilities perform outside demonstrations: with imperfect documents, changing interfaces, and decisions where an error can have real consequences.

Gemini 3 Pro improves AI vision | neversleep.ai