← All posts / Models

First Hypothesis in 100ms: Microsoft's MAI-Transcribe-2-Streaming Debuts at No. 1, Plus Two New Voice Models

Microsoft AI ships its first streaming transcription model with a 2.5% WER, 60 languages, and $0.54/hr pricing, flanked by MAI-Voice-2.1 and a 150ms Flash variant aimed squarely at the voice-agent loop.

First Hypothesis in 100ms: Microsoft's MAI-Transcribe-2-Streaming Debuts at No. 1, Plus Two New Voice Models

The voice race has a new scoreboard leader. On October 1, 2026, Microsoft AI launched three speech models in one shot: MAI-Transcribe-2-Streaming, the company’s first streaming transcription model, plus two text-to-speech siblings, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Together they form a complete pipeline for what has quietly become the most competitive arena in applied AI — real-time voice agents that can listen, understand, decide, and speak inside the window where a human still experiences the exchange as a conversation.

A transcription model that doesn’t wait for you to finish

The headline claim is blunt: MAI-Transcribe-2-Streaming debuts at No. 1 on Artificial Analysis for accuracy on both final and partial transcripts. The leaderboard chart, citing the streaming index dated September 28, 2026, puts the model at a 2.5% final word-error rate, a 2.8% first-partial error rate, and 0.13 seconds to final transcription. Microsoft also says the model sits on the Pareto frontier of the benchmark’s accuracy-versus-latency evaluation — a pointed way of saying that the usual trade-off, where accuracy costs you speed, doesn’t apply here.

The mechanics matter more than the medal. Traditional speech-to-text waits for a utterance to end before returning anything usable. MAI-Transcribe-2-Streaming instead emits its first hypotheses — “partials” in the jargon — in just over 100 milliseconds of receiving audio, then continuously revises them as context accumulates before committing a stable transcript. That changes what applications can do: a voice agent can begin reasoning or calling tools while the user is still mid-sentence, and live captions can render as people talk rather than a beat behind.

The model covers 60 languages with automatic, continuous language detection — it can follow a speaker who code-switches without restarting a session, a scenario that breaks many single-language ASR systems. Microsoft claims that in internal evaluations, words appear in the transcript twice as fast as its closest competitor for real-time dictation and subtitling.

Worth noting what the benchmark actually measures. Artificial Analysis’s AA-WER Streaming index evaluates models where audio streams in chunk by chunk, drawing on roughly eight hours of audio from three datasets — AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%) — a mix chosen for diverse accents, domain-specific vocabulary, and hostile acoustic conditions. Time-to-final and time-to-first-partial are measured from the end of speech detected by the SileroVAD voice-activity detector.

The two voice models

MAI-Voice-2.1 is Microsoft’s strongest multilingual TTS yet, spanning 23 languages and 26 locales. Its signature trick is cross-language speaker identity: one voice speaks all supported languages with a genuinely native accent in each, rather than dragging a single accent across language boundaries. Microsoft’s example is a tutoring app that switches languages mid-lesson without swapping teachers, keeping the brand’s voice intact everywhere. It’s priced at $22 per 1M characters.

MAI-Voice-2.1-Flash targets high-volume, latency-sensitive workloads. It generates up to 45 seconds of audio with 150ms end-to-end latency, delivers 55% faster model inference, and is roughly 60% cheaper than comparable models at $15 per 1M characters. Microsoft explicitly positions Flash as the natural partner for MAI-Transcribe-2-Streaming — the pair brackets the voice-agent loop, buying back latency on both the listening end and the speaking end.

Both voice models support cloning across all supported languages from a few seconds of reference audio, with built-in consent guardrails that Microsoft says prevent misuse — a necessary control in a year when voice-cloning abuse has drawn regulatory attention in several jurisdictions.

One striking data point: in a 4,000-listener Turing test, 50.3% of listeners rated MAI-Voice as equally or more human-like than actual human recordings. Crossing the 50% line in a forced comparison against real speech is the kind of milestone that moves TTS from “impressive demo” to “deployment default.”

Context: the MAI audio stack matures fast

This launch extends an audio line that already had momentum. MAI-Transcribe-2, the earlier non-streaming recognition model, was billed by Microsoft as the fastest, most accurate, and cheapest in the world when it shipped. The new streaming variant closes the one gap that mattered for live agents. More broadly, the MAI model family — Thinking-1 for reasoning, Code-1.1-Flash for engineering, Image-2.6 for generation — shows Microsoft building a full-stack, first-party alternative to relying solely on partner models.

The distribution strategy is equally telling. All three models are available through Microsoft Foundry, the MAI Playground, Vercel, and Azure Voice Live, with the voice models also on OpenRouter and LiveKit listed as coming soon. Microsoft also shipped Chatter, a live demo in the MAI Playground where users can talk to an assistant powered by the new transcription and voice models.

Why it matters

Voice is where AI latency budgets are most unforgiving. A text agent that takes three seconds to reply is fine; a voice agent that takes three seconds feels broken. The economics are equally sharp: at $0.54 per hour of audio (introductory pricing through year-end) for transcription and $15 per 1M characters for speech, Microsoft is pricing the full hear-understand-speak loop within reach of consumer-scale applications, not just enterprise pilots.

The competitive read: Google shipped its own transcription push weeks earlier with Gemini 3.5 Transcribe, and OpenAI, ElevenLabs, and Decagon’s Chord are all fighting over the same voice-agent turf. Microsoft’s bet is that winning the leaderboard on a neutral third-party benchmark — and pairing it with aggressive pricing and broad distribution — is the fastest way to become the default plumbing for the voice-agent wave. The No. 1 spot will be contested monthly. The 100ms-first-hypothesis approach to streaming is the part competitors will have to match.