AI News
AI News AgentModel releaseX.ai2 min read

Grok-1.5 improves reasoning and processes more context

xAI introduces Grok-1.5, a model with better results in mathematics and programming and a context window of up to 128,000 tokens. Its rollout will be gradual for evaluators and current Grok users on 𝕏.

xAI has introduced Grok-1.5, a new version of its artificial intelligence model with better results in mathematics, programming, and long-text comprehension. The company plans to offer it first to current Grok users and a group of evaluators in the days following the announcement.

More capacity for solving problems

The clearest improvement appears in reasoning tasks. In tests published by xAI, Grok-1.5 scored:

  • 50.6% on MATH, an evaluation of school and pre-university-level math problems.
  • 90% on GSM8K, another test focused on math problems.
  • 74.1% on HumanEval, which measures the ability to generate code and solve programming problems.

These are results from specific tests, not a guarantee that the model will always get the answer right. Model comparisons also need to be read carefully because each evaluation may use a different number of examples, instructions, and conditions.

In xAI's table, Grok-1.5 ranks above Grok-1 across all of these measurements. On MATH, for example, it rises from 23.9% to 50.6%. On GSM8K, it goes from 62.9% to 90%.

It can read much longer documents

Grok-1.5 expands its context window to 128,000 tokens. Context is the amount of text a model can take into account during a conversation or task. According to xAI, this multiplies the available length by 16 compared with the previous version.

In practice, this makes it possible to work with long documents without splitting them into many parts. For example, you could ask it to compare several chapters of a report, find a specific clause in a long contract, or review a code project with many files.

xAI also says that Grok-1.5 correctly retrieved information hidden in texts of up to 128,000 tokens during the Needle In A Haystack test. This evaluation measures whether the model can locate a specific piece of information within a very large volume of text.

The rollout will be gradual

Grok-1.5 will not be available to everyone immediately. xAI says it will first reach early evaluators and users who already have access to Grok on the 𝕏 platform. The company then plans to expand availability and add new features.

The model was trained using proprietary infrastructure based on JAX, Rust, and Kubernetes. Among other things, the system can detect problematic servers and automatically remove them during training, while also restarting the process with less lost time.

For you, the most practical news is not just that Grok achieves higher scores. It is that the model can keep more information active while it works. That could lead to better results when analyzing long documents, writing or reviewing code, and solving multistep tasks, although its performance in daily use will need to be tested once more people have access.

Grok-1.5 points to the direction of competition between AI models: not just answering questions, but working with more information and sustaining complex tasks for longer. What matters now is whether its laboratory improvements hold up in real conversations and work.

Grok-1.5 improves reasoning and processes more context | neversleep.ai