toolcall.
ModelsJul 29, 2026, 12:00 UTCUpdated Aug 11, 2026

Introducing Grok Voice Think Fast 2.0

xAI unveils its most advanced speech-to-speech voice model

xAI has introduced Grok Voice Think Fast 2.0, its latest speech-to-speech voice model designed to improve intelligence, transcription accuracy, and conversational ability. This model demonstrates significant advancements over Grok Voice Think Fast 1.0 and competing models in several benchmarks. It achieves an overall speech-to-speech quality index of 82.9%, outperforming others like GPT-Realtime-2.1 and Gemini 3.1 Flash. The model excels in speech reasoning, conversational dynamics, and agentic performance, with a notably faster time to first audio at 0.70 seconds. Transcription accuracy is a key strength, showing a 1.5 to 2 times improvement over dedicated transcription models such as Deepgram Nova 3 and ElevenLabs Scribe v2, especially in noisy settings where the gap widens to about 10 times. Grok Voice Think Fast 2.0 also reasons while speaking, making tool calls more responsive without added latency. Its conversational style is refined through reinforcement learning to produce shorter sentences and clearer interactions. The model is compatible with existing prompts and is priced at $0.08 per minute of audio, with no required action for users to upgrade from the previous version.

Sources

Mentioned

grokxai