xAI (now SpaceXAI) shipped grok-voice-think-fast-2.0, a native speech-to-speech voice model. Audio in, audio out, no text pipeline in between. The trick: it reasons in parallel with speaking. Voice models normally face a brutal tradeoff — think first and users hear dead air, skip thinking and get dumb answers. This one does both at once, with zero latency cost.
The numbers against OpenAI and Google
On Artificial Analysis’ speech-to-speech benchmark: 82.9%, up from 75.7% for v1.0 — ahead of GPT-Realtime-2.1 (79.1%) and Gemini 3.1 Flash (69.5%). Time to first audio dropped from 1.25s to 0.70s. Transcription accuracy is 1.4x better across 24 languages, using roughly 60% fewer reasoning tokens. xAI ran it on Starlink’s phone service and reports higher sales conversion and support containment.
The API switch happens today
$0.08 per audio minute through the xAI API. On August 5 the grok-voice-latest alias auto-routes to 2.0 — pin v1.0 manually if you don’t want the upgrade. Agent Builder handles the obvious use cases: phone support, sales calls, voice agents. Same release adds grok-imagine-video-1.5 with native 1080p text-to-video. The real-time voice race just became a reasoning race.
You Might Also Like
- Openai gpt Realtime 2 Translate Whisper Three Voice Models one api Several Startups Erased
- Openai gpt Realtime 2 1 gpt Realtime 2 1 Mini cut Voice Agent Latency by 25
- Openai Realtime Voice Webrtc Stack the Infra Blueprint Every Voice Agent Startup now has to Compete With
- Hugging Face Speech to Speech Open Source Local Voice Agents is the Openai Realtime Clone you can run on Your own gpu
- Assemblyais Voice Agent api Undercuts Vapi and Retell With 307ms stt Latency

Leave a comment