xAI launches Grok 4.1 with better conversation and accuracy
xAI has made Grok 4.1 available to all users on the web, X, and its mobile apps. The update improves conversation, creative writing, and emotional interaction, while aiming to reduce errors in informational responses. According to xAI, users preferred this version in 64.78% of blind comparisons with the previous Grok.

Grok 4.1 is now available to all users on Grok.com, X, and the iOS and Android apps. The model is immediately available in automatic mode and can also be selected manually from the model picker.
The update is not focused only on answering harder questions. xAI has worked specifically on how Grok converses: interpreting intent more accurately, maintaining a more consistent personality, and offering more useful responses in creative, emotional, and collaborative interactions.
What changes compared with the previous Grok
For two weeks, xAI tested preliminary versions of Grok 4.1 on a growing share of real traffic across its platforms. The responses were compared blindly, without evaluators knowing which version had generated each one.
According to the company, users preferred Grok 4.1 over the previous model in 64.78% of comparisons. The evaluation used real traffic from Grok.com, X, and the mobile apps.
The new version also holds prominent positions on LMArena, a public ranking that compares models through human evaluations. In reasoning mode, Grok 4.1 Thinking, whose internal name is quasarflux, ranks first in Text Arena with 1483 Elo points, a measure of relative performance.
The non-reasoning model, designed to respond faster, reaches 1465 Elo points and ranks second. xAI says this version outperforms other models that do use full reasoning. For reference, Grok 4 ranks 33rd in that classification.
More attention to conversations
xAI also evaluated Grok 4.1 on EQ-Bench3, a test that measures capabilities such as emotional understanding, empathy, perception, and interpersonal interaction. The test uses 45 role-play situations, many of them developed over three conversation turns.
The company also tested the model on Creative Writing v3, a creative writing benchmark with 32 prompts and three iterations per prompt. These tests try to measure something that does not always appear in mathematics or programming rankings: whether a response feels appropriate, coherent, and natural to the person reading it.
To achieve this, xAI applied its reinforcement learning infrastructure to areas such as style, personality, usefulness, and alignment. It also developed methods for other advanced models to automatically evaluate Grok's responses and help improve the system at scale.
Fewer errors in informational questions
Grok 4.1 also includes adjustments to reduce hallucinations, meaning responses that present false information as fact. xAI says it observed a significant reduction in informational queries taken from its production traffic.
The company measured this aspect using a stratified sample of real questions and FActScore, a public test made up of 500 biographical questions. The announcement does not provide a single reduction figure, so the improvement should be understood as an advance observed by xAI, not a guarantee of accuracy.
For you, the most visible change should come in everyday use: less rigid conversations, better responses to ambiguous or emotional messages, and fewer errors when you ask for information. Even so, Grok 4.1 still needs verification when its answer affects important decisions.
The next signal to watch is whether these improvements hold up outside internal evaluations and in large-scale real-world use. Access is no longer the limit. Now the test is how much the experience improves when you use Grok to search, write, or have a conversation.