Launch Overview
On Friday, Grok introduced Voice Transcribe 2.0, a speech‑to‑text model that delivers twice the accuracy of its predecessor while remaining at the same price point. The company states the new model ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
Technical Foundations
Voice Transcribe 2.0 builds on the audio foundation model that powers Grok Voice, which currently handles tens of thousands of customer‑support calls each day, transcribes millions of hours of video narration, and powers voice agents in physical products such as the Grok assistant embedded in Tesla (NASDAQ:TSLA) vehicles.
Accuracy Improvements
In telephony tests using 8 kHz audio, the model achieved a word error rate (WER) of 7.1%, down from 10.6% for the prior version. For conversational audio, WER dropped from 8.7% to 3.3%. Errors in transcribing spoken credentials such as phone numbers and email addresses fell from 7.2% to 3.2%. The model was evaluated against seven competitors across four categories: telephony audio from customer‑support calls, conversations with Grok, spoken credentials, and short multilingual voice commands.
Language Support
Voice Transcribe 2.0 supports dozens of languages with automatic detection and can handle mid‑recording language switches. For short phrases in 19 languages, the WER decreased from 20.6% to 6.8%.
Adoption and Pricing
Atlassian’s Loom platform has adopted the new model for transcribing all its video recordings. "We’ve always believed the best way to move work forward is to capture context once and let it flow everywhere," said Sanchan Saxena, SVP of Teamwork Collection at Atlassian. Pricing remains unchanged at $0.10 per hour for batch transcription and $0.20 per hour for streaming, with speaker diarization, word‑level timestamps, and key‑term biasing included at no extra cost.
Implementation Timeline
The model will become the default offering in Grok’s Speech‑to‑Text API in the coming weeks.