← All posts / Models

3x Better Transcription for the Same Price: SpaceXAI's Grok Voice Transcribe 2.0 Takes On the Speech AI Market

SpaceXAI cuts word error rate from 20.6% to 6.8% in one generation while holding prices at $0.10 per batch hour — no code changes required, and Loom is already on board.

3x Better Transcription for the Same Price: SpaceXAI's Grok Voice Transcribe 2.0 Takes On the Speech AI Market

Earlier this week, SpaceXAI quietly shipped one of the most aggressive speech-to-text upgrades of the year. Grok Voice Transcribe 2.0, now live in the company’s Speech-to-Text API, cuts the model’s word error rate from 20.6% to 6.8% on SpaceXAI’s internal multilingual benchmark — a roughly 3x accuracy improvement in a single generational leap — while keeping prices exactly where they were: $0.10 per audio hour for batch transcription and $0.20 per hour for streaming.

In a market where the major speech vendors treat every accuracy tier as a pricing tier, “same price, three times better” is an unusual posture. It signals that xAI’s voice stack, like its text stack, intends to compete primarily on price-performance — and it arrives just as the voice-agent economy is scaling from demos to production traffic.

What shipped

Grok Voice Transcribe 2.0 is available immediately for both recorded audio and live streams. According to SpaceXAI’s release notes, version 1.0 remains the default model for now, but version 2.0 is expected to take over “soon,” with 1.0 headed for deprecation in the following weeks. Customers who need the older version during the transition can pin it; existing integrations require no code changes — developers simply pass grok-voice-transcribe-2.0 as the model name.

The feature set reads like a checklist of everything production transcription workloads actually need:

  • Speaker diarization — automatic speaker labels woven through the transcript
  • Word-level timestamps — alignment precise enough for captioning and media search
  • Confidence scores — per-word signals for downstream quality filtering
  • Up to eight audio channels — multi-mic recordings transcribed in a single call
  • Key-term biasing — domain vocabulary (product names, jargon) boosted on request
  • Formatting intelligence — numbers, currency, and email addresses rendered correctly instead of phonetically
  • Filler-word removal — “um” and “uh” scrubbed for cleaner readouts

SpaceXAI also claims the model holds up on the audio that actually breaks transcription systems: heavy accents, degraded phone lines, overlapping speakers, and mid-sentence language switching across dozens of languages.

The numbers, with an asterisk

The headline 20.6% → 6.8% WER improvement comes from SpaceXAI’s own internal multilingual test suite, which means it’s company-reported rather than independently verified. The practical question for any team evaluating the model is whether those gains hold on the audio their product actually receives — call-center noise, meeting-room reverb, field recordings — rather than on curated benchmarks.

SpaceXAI points to one external data point: on Artificial Analysis’s leaderboard, the company says Transcribe 2.0 ranks first among 32 streaming models. That’s a meaningful independent-ish signal, though leaderboard audio distributions and production traffic are different beasts.

The anchor customer is Atlassian’s Loom, which is already running every uploaded video through the new model. That’s a serious workload — Loom processes a large volume of async video messages daily — but one happy customer doesn’t establish performance across the wild variety of workloads the market serves.

Pricing warfare

The pricing is where this release gets strategically interesting. At $0.10 per batch hour and $0.20 per streaming hour, Grok Voice Transcribe 2.0 undercuts or matches essentially every major competitor:

ProviderBatch (per hour)Streaming (per hour)
Grok Voice Transcribe 2.0$0.10$0.20
AssemblyAI Universal-2$0.15~$0.27+
ElevenLabs Scribe$0.22$0.39
Deepgram~$0.26~$0.31+
OpenAI (Whisper-class)~$0.36 ($0.006/min)—

The speech AI market has been drifting toward commoditization for two years, but this is one of the clearest “accuracy up, price flat” moves any major lab has made. If the WER claims hold up under third-party testing, the effect on competitors’ pricing will be immediate — Deepgram, AssemblyAI, and ElevenLabs all justify premium tiers on accuracy claims that a $0.10/hour model now contests directly.

Why now: the voice-agent boom

The timing is not accidental. Voice AI is having its breakout year. Real-time voice agents — customer support lines, drive-thru ordering, appointment scheduling, in-car assistants — all live or die by streaming transcription quality. A streaming model that ranks first among 32 on Artificial Analysis at $0.20/hour is precisely the ingredient voice-agent builders have been asking for, and SpaceXAI knows it. The Grok Voice API family already bundles speech-to-speech, text-to-speech, and speech-to-text with tool calling and real-time search; Transcribe 2.0 makes the transcription leg of that stack competitive with anyone’s.

There’s also a platform play underneath. Transcription output feeds retrieval, summarization, agent memory, and compliance pipelines. Whoever owns the STT layer increasingly owns the data spigot for everything downstream — and at $0.10/hour, the economics favor routing enormous volumes through the winner.

What to watch

Three open questions will decide whether this release is a milestone or a footnote:

  1. Independent benchmarks. Will the 6.8% WER claim survive third-party evaluation on messy, real-world audio? The gap between internal suites and production audio is where speech vendors traditionally get humbled.
  2. The default flip. Version 1.0 is still the default model. When 2.0 takes over and 1.0 is deprecated, any regression in edge cases will surface quickly at scale.
  3. Competitive response. Deepgram and AssemblyAI have both been racing down the cost curve; ElevenLabs has premium brand pull. A price war in speech AI would accelerate voice-agent deployment across every industry that runs phone lines.

For developers, the calculus is simple: the upgrade is free to try, requires no code changes, and pins are available if it disappoints. Expect adoption to move fast — which is exactly what SpaceXAI is counting on.