← All posts / Models

Meta's Muse Voice Transcribe Listens Like a Human: 20+ Speaker Diarization, 70+ Languages, One Hour Sessions

Meta Superintelligence Labs ships its first real-time audio perception model: streaming ASR with adaptive delay trained via RL, native code-switching, and a claimed #1 spot on Artificial Analysis speech-to-text rankings.

Meta's Muse Voice Transcribe Listens Like a Human: 20+ Speaker Diarization, 70+ Languages, One Hour Sessions

Meta Superintelligence Labs (MSL) has released Muse Voice Transcribe, its first real-time audio perception model — and the announcement reads less like a speech-to-text refresh and more like the foundation layer for the “personal superintelligence” vision Mark Zuckerberg laid out earlier this year. The model went live September 1, 2026, and it is available today through the Meta Model API, Meta AI for Mac, and Muse Code, with one-click voice dictation already switched over to it across Meta’s apps. Just hold the Fn key to try it on Mac.

What it does

On paper, the capability list is a checklist of everything that has historically been hard in streaming audio:

  • Streaming ASR — real-time speech-to-text, not batch post-processing
  • Diarization with 20+ speakers — attributing “who said what” live, in a single pass
  • Endpointing — detecting when a user starts and finishes speaking, the make-or-break skill for voice agents
  • Multilingual input with seamless code-switching — trained on more than 70 languages, 25 of them extensively verified at launch
  • Long context — audio sessions exceeding one hour, natively, with no required post-processing
  • Context biasing — language, keyword, and context hints that teach the model your contacts, jargon, and place names

Meta claims the model ranks first on Artificial Analysis’s streaming speech-to-text leaderboard and on public diarization benchmarks (rankings as of September 1, 2026). The Information’s briefing highlighted the headline numbers: more than 70 training languages and hour-long sessions with over 20 speakers. Artificial Analysis confirmed the 70+ language training set and 25 extensively verified languages independently.

The architecture: audio as a token stream

The technically interesting part is how MSL framed the problem. Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family — the same lineage as the agentic models Meta has been shipping to the Meta Model API since July. Audio arrives in 80ms chunks (12.5 Hz), and each chunk is transformed into a single soft token. At every step, the model makes a decision: keep listening, or emit a text token.

That decision is mediated by special tokens. When the model wants more audio context before committing to a word, it predicts a <|next_audio|> token, which gets replaced by the actual next audio chunk. When the stream stops, a <|empty_audio|> token signals that no more audio is coming, and the model flushes its remaining text predictions.

This “delay” is the crux of streaming transcription. The longer the model waits, the more accurate the transcript — and the worse the latency. Muse Voice Transcribe’s answer is adaptive delay: the model dynamically changes how long it waits for each word based on difficulty. Meta trained this behavior with reinforcement learning, combining a word-error-rate (WER) reward and a delay reward multiplicatively. The result, Meta says, sits on the Pareto front of the speed-accuracy trade-off as measured by time-to-final-transcription.

Diarization and endpointing fall out for free

Because ASR, diarization, and endpointing all share one autoregressive backbone, Meta added the other two tasks as special tokens on top of the ASR foundation:

  • Diarization introduces <|start_of_turn|> to mark speaker switches and <|speaker_{A-Z}|> tags to identify who is talking. Turn boundaries are predicted as soon as a switch happens; the speaker tag itself is predicted at the end of each chunk. A single speaker’s audio can be split across multiple <|start_of_turn|> segments, all carrying the same tag.
  • Endpointing adds <|speech_onset|> and <|speech_endpoint|> tokens — exactly what a voice assistant needs to know when you’ve finished a request (“Hey Meta, what’s the weather in Menlo Park?”) and it’s safe to respond.

All tasks are trained together with streaming ASR, with extra rewards stacked on the base ASR reward for diarization and endpointing respectively. This is the same “one backbone, many heads” philosophy that has been winning across multimodal research — and it explains how a first-release model can post top diarization numbers without a separate speaker-verification stack.

Code-switching is the killer demo

The demos Meta published are unusually candid about real speech. One features eight people in one room talking over each other about Los Angeles hiking trails — overlaps, interruptions, accents, tangents — tracked live with speaker labels. Another is a one-hour, eleven-speaker conversation transcribed end to end with no post-processing.

For bilingual speakers, the standout is the code-switching demo: a Mandarin-English mix inside single sentences (“昨天我在 local desktop 上用 Ollama 跑了下 Meta 的 Muse Glimmer… 整个 setup 非常 smooth”), dense with technical terms like speculative decoding, 4-bit GGUF quantization, and VRAM. The transcript is clean. Meta’s own researchers argue this is non-negotiable: if a “personal” assistant can’t handle “明天九点有个 doctor appointment” mid-sentence, it isn’t actually for everyone.

The strategic read

Two things make this launch more than a benchmark flex.

First, it’s MSL’s first perception model. Everything Meta has shipped from its superintelligence push so far has been generative or agentic. A personal agent that runs on AI glasses “in real conversations, not just voice commands” — Zuckerberg’s own framing — needs ears before it needs a mouth. Streaming endpointing and diarization are precisely the components that turn a chatbot into something that can sit in a meeting or a family dinner and participate.

Second, the distribution is already live. Voice dictation across Meta AI and Muse Code is powered by Muse Voice Transcribe as of today, and the model carries the Muse Spark family pricing on the Meta Model API — $1.25 per million input tokens and $4.25 per million output tokens, with a cheaper Contributor tier. That undercuts incumbents on a capability (streaming diarization at 20+ speakers) that OpenAI and Google currently serve through separate, stitched-together pipelines, as The New Stack noted in declaring that Meta “just beat OpenAI and Google at real-time transcription.”

The open question is openness. The blog post says nothing about open weights, and “Muse” branding has historically spanned both open and closed releases. For an industry that has watched Meta give away Llama and (reportedly) muse over the same for smaller Muse variants, a closed API-only audio stack would be a notable strategic signal — and for developers building transcription products on top, a risk worth pricing in.

Why it matters

Transcription looks like a commodity until you try to do it live, in a room full of people, in two languages, for an hour. Every one of those constraints has historically forced a trade-off or a pipeline of bolted-together models. If Muse Voice Transcribe’s claims hold up in independent testing, the launch compresses that stack into a single model decision loop — and hands Meta the input layer for the agent era at a moment when OpenAI, Google, and Anthropic are all racing to own the assistant surface.