Ten Cents an Hour: Microsoft's MAI-Transcribe-2 Turns Speech Recognition Into a Utility
Microsoft's MAI-Transcribe-2 tops the FLEURS benchmark across 60 languages at a 5.2% WER and undercuts OpenAI, Google, and ElevenLabs at $0.10 per audio hour — a 72% price cut in five months.
Speech-to-text just became a commodity, and Microsoft is the one holding the pricing gun. This week the Microsoft AI team shipped MAI-Transcribe-2, an in-house speech recognition model the company bills as “the fastest, most accurate and cheapest speech recognition model in the world.” The headline number is hard to look away from: an introductory price of $0.10 per audio hour, available through the end of 2026.
That is not a marginal discount. When the first MAI-Transcribe shipped in April 2026, it cost $0.36 per hour. Five months later, the successor undercuts its own predecessor by roughly 72% — while adding languages, features, and benchmark wins. In a market where call-center platforms, media companies, and meeting-note startups routinely process six-figure volumes of audio, a drop from $36,000 to $10,000 on 100,000 hours of audio is large enough to force every procurement team to rebid.
What the model actually does
MAI-Transcribe-2 covers 60 languages, up from 43 in the June 1.5 release. On the FLEURS multilingual benchmark — the standard cross-lingual speech recognition test — Microsoft says the model ranks first across all 60 languages with an average word-error-rate of 5.2%, beating Gemini 3.5 Transcribe, OpenAI’s GPT-Transcribe, Whisper V3-Large, and ElevenLabs Scribe v2.
Speed is the second axis of the attack. On Artificial Analysis measurements cited by Microsoft, MAI-Transcribe-2 runs 10× faster than GPT-Transcribe, 7× faster than ElevenLabs Scribe v2, and 5× faster than Gemini 3.5 Transcribe, while defining what the company calls the Pareto Frontier for accuracy versus latency — no competing model is simultaneously faster and more accurate. It ranks second on the Artificial Analysis word-error-rate leaderboard overall, which the company frames as continued “hill-climbing” from earlier versions.
The feature list reads like a checklist of everything enterprise buyers have spent the last two years asking for:
- Speaker diarization — recordings are segmented by speaker, with each segment attributed to the right person and tagged with offset and duration metadata.
- Word-level timestamps — every word carries precise timing, enabling subtitle alignment, in-transcript search, and audio navigation.
- Keyword biasing — supply a phrase list of product names, acronyms, or domain jargon, and recognition leans toward them (as hints, not forced output).
- Configurable styles —
verbatimmode preserves filler words, false starts, and self-corrections for compliance and QA workloads;cleanmode strips them for readable notes and captions. - Code switching — blended-language speech like Hinglish and Spanglish is handled mid-utterance rather than forcing one language per recording.
- Automatic language identification, plus robustness to background noise, overlapping speech, and variable microphone quality.
Under the hood of the API
Technically, MAI-Transcribe-2 is available in public preview through the Azure Speech service’s Fast Transcription API. Developers enable it by setting enhancedMode.enabled to true and enhancedMode.model to "MAI-Transcribe-2" on the speechtotext/transcriptions:transcribe REST endpoint (api-version 2025-10-15). Inputs are audio files under 300 MB in WAV, MP3, or FLAC — which means the model currently targets batch workloads like captioning, clinical notes, call-center documentation, and content creation rather than real-time streaming. Beyond Azure, Microsoft is pushing the model through Foundry, MAI Playground, and notably OpenRouter, putting its first-party MAI family in direct price competition with the OpenAI models it also resells inside Azure.
The preview caveat matters: no SLA yet, and Microsoft explicitly warns against production workloads until it generalizes. But the pricing signal is already out, and competitors have to respond to it now.
Why the price collapse is the real story
Speech recognition occupies a unusual position in the AI stack: it is one of the few markets where buyers can compare price, latency, language coverage, and accuracy on a spreadsheet. There is no vibes-based evaluation layer, no subjective taste question like with image generation or writing. When a vendor publishes benchmark wins and a 3–10× price advantage simultaneously, the rational response is to switch.
That transparency is precisely why Microsoft chose this battleground. The MAI family — Transcribe, Voice, Image, and the reasoning models Microsoft has been shipping to Foundry — is Redmond’s bid to be seen as a first-party model developer rather than OpenAI’s distribution channel. Every workload that moves from GPT-Transcribe to MAI-Transcribe-2 marginally rebalances that relationship.
The deeper implication is what commoditization does to the startup ecosystem. Whisper-class quality at ten cents an hour erases the “better transcription” pitch entirely. What survives is workflow (editorial integration, compliance retention, domain fine-tuning), regulatory positioning (healthcare and legal certification), and proprietary domain data. Raw transcription quality is no longer a moat — it is table stakes delivered by a hyperscaler at utility pricing.
There is also a training-data flywheel argument. Call centers, meeting recorders, and captioning pipelines generate billions of hours of audio daily. The vendor that prices low enough to capture that flow gains the largest real-world speech corpus on earth — noisy, multilingual, code-switched audio that clean academic benchmarks never capture. Microsoft’s aggressive pricing is not just market-share strategy; it is data-acquisition strategy, and it compounds.
The competitive ledger
For OpenAI, GPT-Transcribe now costs roughly an order of magnitude more than MAI-Transcribe-2 while running ten times slower on third-party measurements — an untenable position in a spreadsheet-decided market. For Google, Gemini 3.5 Transcribe retains a quality-perception edge but faces a 5× latency gap. For ElevenLabs, whose Scribe v2 was the startup-standard for accuracy, a hyperscaler undercutting on price at equal-or-better WER is the classic squeeze: their escape route is voice generation and agentic speech products, not transcription alone.
Microsoft’s own economics are helped by the fact that inference runs on infrastructure it owns end to end — Azure regions worldwide, its own silicon strategy, and a Speech service that has existed for a decade. Ten cents an hour is only catastrophic pricing if you are renting someone else’s GPUs to deliver it.
The move also lands amid a broader pattern this week: as frontier chat models grab headlines with AGI claims, the sharpest competition has quietly shifted to the boring, high-volume, metered workloads — transcription, translation, captioning — where buyers actually measure things. Whoever wins those utility tiers owns the default rails that every agent, robot, and meeting assistant will need to speak through.
MAI-Transcribe-2 is in public preview now on Azure Speech, Foundry, and OpenRouter at the $0.10/hour introductory rate through the end of 2026. The five-month journey from $0.36 to $0.10 tells you where this curve goes next: toward zero, fast.
Sources
- [1] https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/
- [2] https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe
- [3] https://venturebeat.com/infrastructure/microsoft-ais-mai-transcribe-2-undercuts-openai-google-and-elevenlabs-on-price-and-speed
- [4] https://www.unite.ai/mai-transcribe-2-tops-fleurs-benchmark-across-60-languages-microsoft-says/
- [5] https://techstartups.com/2026/09/04/top-tech-news-today-september-4-2026-amazon-google-microsoft-nvidia-openai-tesla-more/