AI News
AI News AgentModel releaseX.ai3 min read

xAI launches Grok-2 in beta with improved reasoning

xAI introduces Grok-2 and Grok-2 mini in beta for 𝕏 Premium and Premium+ subscribers. The company says Grok-2 outscored Claude 3.5 Sonnet and GPT-4-Turbo in the LMSYS leaderboard while preparing an enterprise API for developers.

xAI introduced Grok-2, a new version of its artificial intelligence assistant, along with Grok-2 mini, a smaller model designed to balance speed and quality. Both are initially arriving in beta on 𝕏 for Premium and Premium+ subscribers.

The company says an early version of Grok-2, tested in the LMSYS leaderboard under the name sus-column-r, outscored Claude 3.5 Sonnet and GPT-4-Turbo on Elo rating. This metric summarizes the results of head-to-head matches between models in Chatbot Arena, a platform where users compare responses without knowing which system generated them.

The result is xAI's claim about one specific test, not a general victory across every task. In the evaluation table published by the company, Grok-2 ranks above GPT-4o in some tests but below Claude 3.5 Sonnet in others, including advanced scientific knowledge and programming.

What improves in Grok-2

Grok-2 is designed to answer questions, write, program and solve problems more accurately than Grok-1.5. xAI also highlights improvements in tool use and in analyzing retrieved information. That means identifying which data is needed, organizing a sequence of facts and discarding irrelevant posts.

In the company's academic tests, Grok-2 achieved these results:

  • 56% on GPQA, a test of advanced scientific questions, compared with 48% for GPT-4-Turbo and 59.6% for Claude 3.5 Sonnet.
  • 87.5% on MMLU, which measures general knowledge, compared with 88.7% for GPT-4o.
  • 76.1% on MATH, focused on mathematical problems, compared with 76.6% for GPT-4o.
  • 88.4% on HumanEval, a code-generation test, compared with 92% for Claude 3.5 Sonnet.
  • 69% on MathVista, which combines images and mathematics, above GPT-4o at 63.8%.
  • 93.6% on DocVQA, a test of questions about visual documents, compared with 95.2% for Claude 3.5 Sonnet.

Grok-2 mini generally delivers lower results but requires fewer resources. In practice, that can mean faster responses and lower costs when it is used for simple tasks such as summarizing texts, drafting messages or classifying information.

What changes for people who use 𝕏

Grok-2 and Grok-2 mini are being added to the Grok tab in the 𝕏 app. Premium and Premium+ users can try them after updating the app. The assistant combines its text and vision capabilities with recent information published on the platform.

That makes it useful for tasks such as these:

  • Summarizing or analyzing a post and its replies.
  • Finding context about a topic circulating on 𝕏.
  • Helping draft a reply.
  • Answering questions about an image or document.
  • Supporting programming and writing tasks.

xAI is also testing image generation with FLUX.1, a model developed by Black Forest Labs. The company also anticipates a future multimodal understanding feature that will allow users to work more seamlessly with text, images and other types of content.

The API comes later

Developers will be able to access Grok-2 and Grok-2 mini through xAI's enterprise API, whose launch was planned for late August 2024. The platform will include deployments across multiple regions to reduce latency, along with multifactor authentication, traffic statistics, billing analytics and tools for managing teams and users.

For you, the difference will depend on how the model reaches the final product. On 𝕏, access is limited to certain paid plans and offered as a beta, so its responses and features may still change. For businesses, the API opens the door to integrating Grok into internal search tools, customer support assistants or document analysis systems.

The key question now is whether Grok-2's improvements hold up outside controlled tests. The next thing to watch will be its performance in everyday use, especially when working with recent information from 𝕏, images and tasks where a factual error can have consequences.

xAI launches Grok-2 in beta with improved reasoning | neversleep.ai