AI News
AI News AgentModel releaseX.ai3 min read

xAI Introduces Grok 3 with Reasoning Models

xAI introduces Grok 3 and Grok 3 mini, two beta models that can spend more time reviewing their answers and solving complex problems. The family includes a one-million-token context window and arrives alongside DeepSearch, an agent for researching and synthesizing information.

xAI introduces Grok 3, a new family of artificial intelligence models that can spend anywhere from a few seconds to several minutes solving a problem before responding. The company is also launching Grok 3 mini, a more affordable version focused on technical reasoning tasks.

Both are arriving in beta and are still being trained. xAI says they were developed on its Colossus supercomputer with ten times more computing power than its previous models, and that they perform better in mathematics, programming, general knowledge and instruction following.

Two ways to use Grok 3

The main new feature is the Grok 3 (Think) and Grok 3 mini (Think) models. When you activate the Think button, the system spends more time reviewing the problem, testing alternatives, detecting errors and refining its answer.

In practice, that targets tasks such as solving a difficult math problem, debugging a program or analyzing a lengthy document. For simple questions, normal mode responds faster and avoids wasting resources on unnecessary reasoning.

xAI says Grok 3 (Think) achieved these results in its tests:

  • 93.3% on AIME 2025, a US mathematics competition, using its highest level of computation during the response.
  • 84.6% on GPQA, a scientific question benchmark designed for graduate-level knowledge.
  • 79.4% on LiveCodeBench, an evaluation of programming problem generation and solving.

Grok 3 mini also performs well on technical tasks that require less general knowledge: it reached 95.8% on AIME 2024 and 80.4% on LiveCodeBench, according to xAI.

These figures are laboratory results, not a guarantee that the model will always be correct. The models are also still changing as they receive training and user feedback.

One million tokens for long documents

Grok 3 includes a context window of one million tokens, eight times more than xAI's previous models. A token can be a word, part of a word or a symbol, so the figure does not correspond exactly to one million words.

What matters is that the model can work with very long documents, multiple files or complex instructions without losing the initial information as easily. xAI says it achieved 83.3% on LOFT, a benchmark focused on finding and connecting information within lengthy contexts.

The company also reports improvements in image and video understanding, scientific knowledge, general knowledge and code generation. In xAI's published evaluation table, Grok 3 outperforms models such as GPT-4o, Claude 3.5 Sonnet and DeepSeek-V3 on several of these tests, although the results depend on the version, evaluation method and amount of computation used.

DeepSearch aims to go beyond a search

Alongside Grok 3, xAI is introducing DeepSearch, an agent that can consult information online, compare conflicting facts and opinions, and draft a final report. The idea is not to simply display links: it should bring the information together and explain what it means.

For example, it could investigate a recent news story, summarize a scientific topic or compare several sources. Its answers still need to be reviewed, especially when it works with current, controversial or difficult-to-verify information.

Grok 3 is available to X Premium and Premium+ users and is also being rolled out with limits to other Grok users. Premium+ subscribers get immediate access to Think and DeepSearch, while the developer API and enterprise access will arrive in the following weeks, according to xAI.

The most relevant change is not simply that Grok 3 responds better, but that it is beginning to work as a tool that decides when it needs to research, use code or spend more time on a problem. The next test will be whether that reasoning holds up outside exams and proves reliable for real-world tasks, where data changes and mistakes have consequences.

xAI Introduces Grok 3 with Reasoning Models | neversleep.ai