xAI introduces Grok Voice Think Fast 2.0
xAI introduces Grok Voice Think Fast 2.0, a voice model with lower latency, better transcription, and greater ability to use tools during a conversation. The updated version will arrive automatically in `grok-voice-latest` on August 5, 2026, and will cost $0.08 per minute.

xAI has introduced Grok Voice Think Fast 2.0, a voice model designed to respond more accurately, hold more natural conversations, and execute tools with less delay. The update will arrive automatically for many API users on August 5, 2026.
According to Artificial Analysis tests cited by xAI, the model achieves 82.9% in overall voice-to-voice quality, compared with 75.7% for the previous version. It also outperforms GPT-Realtime-2.1 and Gemini 3.1 Flash on this specific index.
What improves in practice
The main difference is that Grok can reason while it speaks. Instead of waiting to finish all processing before starting a response, it analyzes the request in parallel with voice generation. xAI says this improves its ability to handle requests without increasing latency.
Time to first audio drops from 1.25 seconds to 0.70 seconds compared with the previous version. If you use a phone assistant, you will notice the difference as fewer silences after asking a question.
The model also performs better in conversations with interruptions and simultaneous responses, known as full-duplex conversation. In that test, Grok Voice Think Fast 2.0 scores 95.1%, very close to GPT-Realtime-2.1 at 95.7% and above Grok 1.0 at 77.8%.
Its results on tasks that require taking action through tools also improve: the model reaches 56.5% on the τ-voice benchmark, compared with 52.1% previously. xAI says tool calls are often executed before the agent finishes its first sentence.
Better transcription, even with noise
xAI says the model outperforms specialized transcription models in an evaluation involving thousands of short phrases in 24 languages. According to its measurements, it delivers a 1.5 to 2 times improvement over Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4 times improvement over Grok Voice Think Fast 1.0.
The difference is reportedly greater in difficult situations, such as compressed phone calls or environments with heavy background noise. In those scenarios, xAI claims an approximate 10 times advantage over dedicated speech-to-text models. These are the company’s evaluation results, not a guarantee for every language or conversation.
Shorter, less mechanical conversations
Training has also focused on how the model converses. xAI says the model tends to use shorter sentences, ask one question at a time, and avoid unnecessary explanations.
That can be useful in customer service workflows. The agent can guide you step by step through solving a problem, booking a service, or checking an order without giving you a long speech or asking for several things at once.
The company says the improvements work without changing existing instruction prompts. In A/B tests conducted in Starlink’s customer service and sales operations, xAI observed a significant increase in both sales conversion and the number of queries resolved without human intervention.
Automatic changeover and pricing
The grok-voice-latest identifier will move from grok-voice-think-fast-1.0 to the new version on August 5, 2026. You will not need to modify your integrations to receive the update.
Anyone who needs to keep the previous model will have to manually set grok-voice-think-fast-1.0 before that date. Grok Voice Think Fast 2.0 will cost $0.08 per minute of audio.
The next thing to watch is whether these laboratory improvements hold up in long conversations, with varied accents, and in real-world noisy environments. For end users, the promise is concrete: voice assistants that respond sooner, transcribe more accurately, and complete tasks without sounding so rigid.