SpaceXAI Launches Grok Voice Transcribe 2.0 With Double the Accuracy

Grok Voice Transcribe 2.0. Credit: SpaceXAI
SpaceXAI released Grok Voice Transcribe 2.0 on September 18, pitching roughly double the accuracy of version 1.0 at the same API price, and ranking first among 32 streaming speech-to-text models on the public Artificial Analysis leaderboard.
The model rides on the same audio foundation that powers Grok Voice in customer-support calls, video narration, and the Grok assistant inside Tesla vehicles. SpaceXAI says it was trained on live, noisy, multilingual audio rather than clean studio clips.
- Batch transcription: $0.10 per hour of audio
- Streaming transcription: $0.20 per hour
- Diarization, timestamps, and key terms included
- Artificial Analysis streaming final transcript: 2.7% word error rate at 0.49 seconds after end of speech
- Non-streaming AA-WER: 2.3%, improved from 4.0% on version 1.0
SpaceXAI’s own post also highlights stronger multilingual results, including automatic language detection and mid-recording language switches in a single pass. On an internal short-phrase set meant to mimic voice-assistant commands, it says word error rate fell from 20.6% to 6.8%.
Developers can call the new model through the existing Speech-to-Text API. SpaceXAI says Transcribe 2.0 will become the default soon, with version 1.0 deprecated in the coming weeks. Teams that need to stay on the older model can pin grok-voice-transcribe-1.0.
Atlassian is already using the stack on Loom, with SVP Sanchan Saxena saying more accurate transcripts help push recorded context into coding tools. Pricing does not change with the upgrade, so the pitch is higher accuracy without a new rate card.
Want to see more of our stories on Google?
P.S. — Buying a new Tesla? Click here to save $1,000 USD, while supporting independent news.
Help support us by shopping on Amazon here.
Links in this post are affiliate links, so we earn a tiny commission at no charge to you. Thanks for supporting independent media!