Google Launches Gemini 3.5 Transcribe: Speech-to-Text That Cleans Up Your Words
Google's most precise speech-to-text model yet hits GA with 4.0% streaming WER, 85+ languages, disfluency cleanup, and function calling — landing in Gboard, Chrome, Antigravity, and the Gemini API.
For years, voice dictation has been the technology everyone uses and nobody loves. You speak, the model transcribes — and then you spend twice as long fixing “ums,” broken punctuation, mangled product names, and that order number it turned into a random phrase. Google’s answer arrived Wednesday: Gemini 3.5 Transcribe, the company’s most precise speech-to-text model to date, now generally available to developers and rolling out across Google’s consumer surfaces.
The launch is more consequential than a routine model refresh. Google is betting that the next interface battle won’t be won by better chatbots but by voice that actually works — and it is shipping Transcribe into Gboard, Chrome, Google Antigravity, and the Gemini app on macOS on day one, alongside public previews in the Gemini API and the Gemini Enterprise Agent Platform.
Two APIs, one model family
Gemini 3.5 Transcribe ships as two dedicated endpoints built for very different jobs:
- Real-time streaming (
gemini-3.5-transcribe-live) — continuous, bidirectional streaming with sub-second latency via the Live API, aimed at interactive voice apps, live agents, and real-time captioning. - Pre-recorded audio processing (
gemini-3.5-transcribe) — batch transcription of meetings, call logs, and long recordings via the Interactions API, with speaker attribution and word-level timestamps.
The split matters. Voice agents and live assistants can’t tolerate the latency of a batch pipeline, while post-call analytics cares more about attribution and timestamps than instant turnaround. By offering both from the same underlying model family, Google lets teams keep one vendor — and one consistent accuracy profile — across interactive and offline workloads.
The numbers: 4.0% streaming WER, 70% faster finalization
The headline metrics come from Artificial Analysis’s independent benchmarking: an average Word Error Rate of 4.0% in streaming mode and 2.6% for non-streaming use. On the multilingual FLEURS benchmark across a set of top languages and locales, the model achieves 5.50% WER streaming and 5.04% non-streaming, improving over Chirp 3, Google’s previous-generation transcription model.
Latency is the other big leap. Time to final transcription improves by 70% over Chirp 3 — a difference you can feel in a live conversation, where a three-second lag breaks the illusion of talking to something that understands you. Google also emphasizes robustness in noisy, real-world environments and accurate capture of alphanumeric entities like postal codes and order IDs, historically a graveyard for speech recognition.
Smart transcription: the cleanup layer
What separates Transcribe from classic ASR is what Google calls smart transcription. The model is designed to capture natural speech intent, not just acoustic content:
- Disfluency removal — filler words (“um,” “uh”) are stripped automatically, and self-corrections like “let’s meet Tuesday — no, Wednesday” resolve to the corrected intent.
- Auto-formatting — output arrives as polished, structured text rather than a raw word stream.
- Custom vocabulary — specialized jargon and unique spellings can be supplied as a vocabulary list, and the model adapts transcriptions on the fly.
- Multi-speaker identification — pre-recorded audio gets speaker attribution with timestamps for up to three speakers (3+ is experimental).
- Function calling — the model can delegate complex tasks such as image generation and file analysis to other Gemini models. It’s live today in the Gemini macOS app, where a spoken “summarize this file and generate a header image” actually executes.
The function-calling detail deserves attention. It reframes transcription from a terminal output — text, end of pipeline — into an entry point for agentic workflows. The transcript becomes a command surface.
85+ languages and live language switching
Transcribe automatically detects and transcribes more than 85 languages, handling regional accents and diverse dialects, and it manages live language switches mid-stream — a scenario that breaks most systems trained on single-language audio. For global contact centers and multilingual teams, that’s the difference between a demo and a deployment.
Early feedback bears that out. Vivo, Intellitek Health, and Lingopal are cited by Google highlighting the model’s latency, accuracy, and language breadth, and the Live API ecosystem — Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents — is already wiring Transcribe into their real-time media infrastructure so developers can deploy voice interfaces without building the streaming plumbing themselves.
Shipping across surfaces, not just APIs
The consumer rollout is unusually aggressive for a speech model:
- Gboard on Android gets Rambler, which turns spoken thoughts into well-formatted text with fillers filtered out — plus voice-driven edits, spelling corrections, and style changes.
- Chrome will soon let you talk-to-type in any web field, dictating replies, drafting posts, or prompting Gemini by voice.
- Google Antigravity pairs Transcribe with screen context and chat history (with permission) for pinpoint accuracy on file names, agent thoughts, and active documents.
- Gemini app on macOS combines transcription with screen context to power multi-step workflows entirely by voice.
The strategy is clear: Google is normalizing voice as a first-class input everywhere it already has a cursor, and letting developers replicate the same experience through the API.
The road ahead
Competition in speech is intensifying — OpenAI’s Whisper long defined the open-source baseline, and every major lab now treats audio as core multimodal territory rather than an accessory. Google’s edge here is distribution: a transcription model is only as useful as the surfaces it lives on, and few companies can ship one into a keyboard, a browser, and an IDE agent in the same week.
For developers, the entry points are public previews in the Gemini API via Google AI Studio and Google Antigravity, with enterprise availability through the Gemini Enterprise Agent Platform and, soon, Gemini Enterprise for Customer Experience. If the benchmarks hold in production noise, the era of apologizing for your dictation may finally be ending — and the era of talking to your computer like it’s actually listening is beginning.