← All posts / Models

Two Cents a Minute: xAI's Grok Voice Transcribe 2.0 Doubles Accuracy at the Same Price

xAI's new speech-to-text model ranks first among 32 streaming models on Artificial Analysis, halves word error rates on telephony and multilingual audio, keeps batch pricing at $0.10 per hour — and already transcribes for Atlassian Loom and Tesla's in-car Grok assistant.

Two Cents a Minute: xAI's Grok Voice Transcribe 2.0 Doubles Accuracy at the Same Price

Speech-to-text has quietly become the most contested layer of the AI stack. In the last month alone, Google shipped Gemini 3.5 Transcribe with disfluency cleanup and function calling, and Microsoft’s MAI-Transcribe-2 declared itself “the cheapest speech recognition model in the world” at $0.10 per audio hour. On September 18, xAI entered the fray with a salvo of its own: Grok Voice Transcribe 2.0, a speech-to-text model the company says is twice as accurate as its predecessor — at exactly the same price.

The launch matters because of where it comes from. Grok Voice Transcribe 2.0 is built on the same audio foundation model that powers Grok Voice, xAI’s production voice stack. That model already handles tens of thousands of customer-support calls per day, transcribes millions of hours of video narration, and runs the voice assistant embedded in Tesla vehicles. In other words, this is not a lab model chasing benchmark clout — it is the distillation of a real-time, noisy, multilingual production workload, refined with post-training into a standalone transcription product.

The numbers that matter

On the public Artificial Analysis leaderboard, Grok Voice Transcribe 2.0 ranks first for accuracy among 32 streaming models, a field that includes offerings from OpenAI, Google, ElevenLabs, and Deepgram. Internal evaluations tell a consistent story across four production-derived test sets:

  • Telephony (8 kHz customer-support calls): the largest single improvement, where 8 kHz phone audio punishes models trained on clean studio recordings. xAI says it leads every model tested on this set.
  • Conversational audio: conversations with Grok in English.
  • Credentials: phone numbers, email addresses, and account codes read aloud — the place where hallucinated digits cause real damage.
  • Short phrases: in-car style voice commands across 19 languages, where word error rate drops from 20.6% to 6.8%.

That last number deserves emphasis. Short utterances give a model almost no context to infer language or intent, which is why they historically break multilingual systems. Cutting WER by roughly two-thirds on this set is the single clearest evidence that the underlying audio foundation model — not just the transcription head — got substantially better.

The pattern extends to specific audio classes reported at launch: word error rate on conversational audio fell from 8.7% to 3.3%, and on short phrases from 20.6% to 6.8%, against the previous generation.

Pricing: the two-cent minute

Grok Voice Transcribe 2.0 keeps pricing identical to version 1.0: $0.10 per hour for batch transcription and $0.20 per hour for streaming — roughly two cents per minute and four cents per streaming minute. Speaker diarization, word-level timestamps, and key-term biasing are included at no extra cost.

That price point puts xAI in direct collision with Microsoft, which set $0.10 per hour as its introductory rate for MAI-Transcribe-2 through the end of 2026. The difference is positioning: Microsoft sells transcription as an Azure utility; xAI is selling the byproduct of a voice-agent stack that already exists at Tesla-scale. When your training data is live customer-support traffic and in-car commands, the marginal cost of shipping a transcription API approaches the cost of serving it.

Features built for agents, not just archives

The feature list reads like a checklist for voice-agent developers rather than archive processing:

  • Batch and streaming — transcribe recorded files and URLs, or stream in real time.
  • Word-level timestamps and confidence scores for every word.
  • Speaker diarization at no additional cost — previously a separate line item at most providers.
  • Multichannel transcription for up to 8 independent channels.
  • Key term biasing with up to 100 domain terms per request — product names, medical vocabulary, internal jargon.
  • Text formatting that renders numbers, dates, currencies, phone numbers, and emails in written form.
  • Filler-word removal (“um”, “uh”) and smart turn detection for knowing when a speaker finishes.

Smart turn detection is quietly the most agent-relevant item: latency in voice agents is dominated by knowing when the user has actually stopped talking, and barge-in behavior depends on it. Bundling it into the transcription tier rather than a separate endpoint simplifies the pipeline for anyone building on xAI’s stack.

The Loom deal: context-to-code

The most strategically interesting detail in the announcement is the customer story: Atlassian is using Grok Voice Transcribe 2.0 to power Loom, its screen-recording product, after finding it more accurate than their existing solution. Atlassian’s SVP Sanchan Saxena framed the workflow explicitly: record an action plan in Loom, pipe the transcript into Cursor, and the code updates itself.

That is a striking alliance on two fronts. First, it puts an xAI model inside Atlassian’s product surface — territory where OpenAI and Microsoft have been assumed to have the enterprise relationship locked. Second, it describes a concrete “context-to-code” loop: speak the intent, transcribe it, hand it to a coding agent. Voice becomes the input device for agentic development, and transcription accuracy becomes the ceiling on how much of that intent survives the trip.

Multilingual in a single pass

Grok Voice Transcribe 2.0 transcribes dozens of languages, auto-detects the language, and — notably — follows mid-recording language switches in a single pass. Multilingual accuracy is claimed as its largest improvement over version 1.0, with comparisons against ElevenLabs Scribe v2 and Deepgram Nova-3 on multilingual WER.

Code-switching is where transcription systems most often collapse into wrong-language output. For support operations in multilingual regions — and for products like Grok’s in-car assistant operating across markets — single-pass code-switching is a genuine capability, not a checkbox.

Analysis: the voice layer consolidates

Viewed from altitude, the last five weeks of speech-to-text launches sketch a clear picture: the voice layer is consolidating around a handful of providers who own the whole stack — audio foundation model, voice agent runtime, and transcription API — rather than specialists selling a single component.

Google embeds Transcribe in Gboard, Chrome, and its agent platform. Microsoft prices it like a utility to anchor Azure. xAI’s differentiator is vertical: the same model family that answers support calls and talks to Tesla drivers is now the one you can call over an API for two cents a minute. Everyone is converging on the same bundle: sub-4% WER on conversational audio, diarization included, timestamps included, key-term biasing included, and a price that treats an hour of audio as a rounding error.

For developers, the practical consequence is good news: a first-place streaming accuracy rank, two-cent minutes, and no-code-change upgrades for existing Speech-to-Text API integrations. Grok Voice Transcribe 2.0 becomes the default in xAI’s Speech-to-Text API soon, with version 1.0 deprecated in the coming weeks (pin grok-voice-transcribe-1.0 to stay on it during the transition).

For the industry, the consequence is sharper: transcription has stopped being a feature and become a substrate. The companies that own the training data flowing through live voice agents — support calls, car cabins, video narration — will keep compounding an advantage that clean-corpus competitors cannot buy. The two-cent minute is not the end of the price war; it is the price of admission to the agent era, where every word spoken to a machine has to be captured, timestamped, formatted, and understood before anything else can happen.