xAI launches Grok Voice Think Fast 2.0, its next-generation speech-to-speech model
xAI has released Grok Voice Think Fast 2.0, the successor to its real-time voice model, and will automatically migrate all traffic pointed at the grok-voice-latest alias to the new version starting August 5, 2026.
What's new
According to xAI's own release notes, "grok-voice-think-fast-2.0 is now available with Speech to Speech. grok-voice-latest will route to this model starting August 5, 2026." The model is accessible now via the Speech to Speech API for developers who want to adopt it ahead of the automatic cutover; those who need more time can continue pinning the 1.0 model identifier.
Independent coverage of xAI's launch materials reports concrete performance gains over the prior generation: time to first audio falling from 1.25 seconds to 0.70 seconds, an overall Speech-to-Speech benchmark score rising to 82.9% (from 75.7% for Think Fast 1.0), and an agentic-performance score of 56.5% versus 52.1% previously. Coverage also cites broadened language support — evaluated across roughly two dozen languages — and pricing of $0.08 per minute of audio.
Context
Think Fast 1.0 debuted earlier this year as xAI's answer to the growing field of low-latency conversational voice models, competing with offerings from OpenAI's realtime voice line and third-party speech APIs such as Deepgram and ElevenLabs. The "Think Fast" branding signals xAI's specific focus on minimizing the reasoning overhead that typically slows down voice agents mid-conversation — a persistent gap between text-based chat latency and natural spoken dialogue.
The release lands the same week xAI also pushed Grok 4.5 into general availability in the EU and shipped an adjustable voice-activity-detection threshold for its Speech to Text endpoint, part of a broader cluster of audio and agent-tooling updates to the xAI API in late July.
Why it matters
Voice is becoming a proving ground for agentic AI products: any assistant meant to hold a live phone call, staff a support line, or narrate an app in real time lives or dies on latency and turn-taking, not just raw language quality. A sub-second time-to-first-audio figure, if it holds up in production traffic rather than benchmark conditions, would put xAI's offering ahead of most publicly available speech-to-speech competitors on responsiveness.
The automatic alias migration on August 5 is also notable operationally — it means every developer currently building on grok-voice-latest inherits the new model's behavior, pricing, and characteristics without an explicit opt-in, which raises the practical stakes of the improvement claims for anyone already shipping voice products on xAI's API.
Corroborating sources
- Docs.x
https://docs.x.ai/docs/release-notes
“grok-voice-think-fast-2.0 is now available with Speech to Speech.”
- Testingcatalog
https://www.testingcatalog.com/spacexai-launches-grok-voice-think-fast-2-0-on-agent-builder/