AI News
AI News AgentModel releaseX.ai3 min read

xAI introduces Grok 4.7 for long-running AI tasks

xAI introduces Grok 4.7, a model aimed at programming and professional tasks that can take several hours. It maintains Grok 4.6's price, improves its results on several tests and adds new safety measures for cybersecurity and biology use cases.

xAI has introduced Grok 4.7, an AI model designed to handle programming and professional work that can take hours. The company says it checks its answers more effectively, handles longer contexts and maintains the same price and speed as Grok 4.6.

More time to solve complex problems

Grok 4.7 uses a larger base model than its predecessor. It also received longer reinforcement training, with greater emphasis on difficult tasks that require multiple steps and a lot of time to complete.

In practice, this is aimed at work such as reviewing an entire software project, preparing a report from multiple sources or creating a professional presentation. The point is not just to generate an answer, but to keep working for longer, check what it has done and correct errors.

xAI also trained the model to understand Grok's conversational environment natively. According to the company, this improves its performance on general questions, conversations and knowledge tasks that are not exclusively related to programming.

Uneven but competitive results

On CursorBench 4.0, a test focused on long-running programming tasks, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6. xAI places it among the models with the best price-to-performance ratio in this evaluation.

Comparisons with other systems show that its results depend on the type of work:

  • On DeepSWE v1.1, it scored 71.0%, below the 72.7% achieved by GPT-5.6 Solmax, but above the 70.0% scored by Fable 5.1max.
  • On AA Briefcase, which simulates professional tasks lasting several hours, it reached 1,657 points, nearly matching Fable 5.1max at 1,678.
  • On Terminal-Bench 4.0, it scored 38.0%, compared with 37.3% for GPT-5.6 Solmax, though well below the 57.9% achieved by Fable 5.1max.
  • In legal and electrical engineering work, it outperformed the models compared in the tests cited by xAI.
  • In clinical reasoning, it scored 56.7%, below GPT-5.6 Solmax and Fable 5.1max.

These results come from evaluations presented by xAI. They do not mean that Grok 4.7 is the best option for every use case. Its performance varies depending on the task and the level of effort used.

Price and availability

Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens. Tokens are the units the model uses to process text, so the final cost depends on how much you send it and how much it produces.

That is the same price as Grok 4.6. xAI also offers a faster variant with twice the output speed, but at twice the price.

The model is available through:

  • Cursor
  • Grok Build
  • The Grok API
  • Third-party programming tools
  • Model routers and cloud platforms

For you, the most important change is that long programming and office tasks can run with less manual supervision. You will still need to review the results for legal, medical, financial or production work.

New safety measures

xAI says Grok 4.7 includes a completely new protection architecture. The company claims it is its most resilient model yet against attempts to bypass its limits, known as jailbreaks, and that it is also more accurate when rejecting dangerous requests.

In cybersecurity, it allowed 3.3% of risky requests through on HackerBench v0.3, according to xAI's evaluation, while blocking fewer legitimate tasks. The company also says it scored 62.4% on LatchBio's biosecurity test.

The balance matters: an overly restrictive model can block valid defensive analysis, while an overly permissive one can help carry out harmful actions. xAI has already begun giving some cybersecurity partners invite-only access to investigate its simulated attack and defense capabilities.

Grok 4.7 arrives with a clear goal: improve performance on long-running work without raising the standard usage price. The question now is whether that advantage holds up beyond benchmarks, when companies and developers use it for hours on real projects and its mistakes have consequences.