Grok-1.5V brings multimodal AI to images
xAI has introduced Grok-1.5V, its first model capable of understanding images, documents, charts, and screenshots in addition to text. The system stands out in real-world spatial understanding, although its results vary against GPT-4V, Claude 3, and Gemini 1.5 Pro.

xAI has introduced Grok-1.5V, a version of Grok that can understand images as well as text. The model can analyze documents, diagrams, charts, screenshots, and photographs, although initial access will be limited to early testers and existing Grok users.
The update matters because it lets you interact with information that is not written in a message. You can show it an invoice to extract its details, a chart to identify trends, a scientific diagram to explain it, or a screenshot to find what is going wrong.
A model that can see, but does not win at everything
xAI describes Grok-1.5V as its first multimodal model, meaning a system that combines text and images in the same conversation. The company compared it with models such as GPT-4V, Claude 3, and Gemini 1.5 Pro across several tests, using direct answers and not asking them to show their reasoning step by step.
The results are uneven, which matters when announcements present these models as good at every task:
- In visual math questions, it scored 52.8%, ahead of
GPT-4Vand Claude 3 Opus, but only slightly ahead of Gemini 1.5 Pro. - In diagrams, it reached 88.3%, close to Claude 3 Sonnet and Opus, and ahead of
GPT-4Vand Gemini. - In reading text within images, it scored 78.1%, virtually the same as
GPT-4V. - In charts and documents, it fell behind several competitors, with 76.1% on ChartQA and 85.6% on DocVQA.
The most notable figure comes from RealWorldQA, a test created by xAI to measure whether a model understands spatial relationships in the real world. Grok-1.5V scored 68.7%, compared with 67.5% for Gemini 1.5 Pro, 61.4% for GPT-4V, and lower results for Claude 3.
What RealWorldQA measures
The initial dataset includes more than 700 images, each with a question and a verifiable answer. It includes images taken from vehicles, along with other real-world scenes, and the questions may require recognizing positions, distances, or relationships between objects.
For example, it is not enough to identify that there is a car in a photo. The system must understand whether it is in front of or behind another object, which element is on the left, or which one occupies a specific position. That ability matters for assistants that may one day help with physical tasks, navigation, or interpreting their surroundings.
xAI has released the dataset under the CC BY-ND 4.0 license and offers a 677 MB download. The company says it plans to expand it as its multimodal models improve.
What changes for you
For now, the practical change is limited: Grok-1.5V is not yet available to everyone. It will first reach early testers and existing Grok users, with no specific date announced.
If access expands, Grok will be able to compete on tasks that have so far driven adoption of ChatGPT, Claude, and Gemini: summarizing scanned documents, interpreting charts, explaining technical images, or answering questions about photographs. But the figures themselves show that its performance depends heavily on the type of image.
xAI's next step will be to improve both understanding and content generation across different modalities, including images, audio, and video. What matters is not only whether Grok can recognize an image, but whether it can do so reliably when the answer has practical consequences.