← All posts / Models

Google's Gemini 3.8 Live Thinks While It Speaks: Native Speech-to-Speech Models Claim the Voice Crown

Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking — speech-native models that reason mid-sentence, call tools in the background, and take #1 on the Speech-to-Speech Quality Index.

Google's Gemini 3.8 Live Thinks While It Speaks: Native Speech-to-Speech Models Claim the Voice Crown

Voice was supposed to be the natural interface to AI — until you actually tried to have a complicated conversation with an assistant. Most voice agents today still bolt a text model between a speech recognizer and a synthesizer, which means they either answer instantly and shallowly, or go silent while they “think” and leave you staring at a spinning indicator. On September 15, 2026, Google attacked that trade-off head-on with the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of native speech-to-speech models that the company calls its “most advanced live dialogue models yet.”

Two models, two jobs

The two variants split along a clean line.

Gemini 3.8 Live is the workhorse, “built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding.” It processes visual inputs in near real-time — point your camera at a chessboard and it plays; walk a new hire through an office and it answers live questions from what it sees. Crucially, it executes tools and API calls in the background while the conversation keeps flowing: the model acknowledges your request, keeps chatting, and reports back when the task completes rather than freezing mid-turn.

Gemini 3.8 Live Extended Thinking is the heavy-lifting sibling, “built for high-complexity tasks, with increased intelligence and multi-step reasoning.” Its signature trick is that it reasons and speaks simultaneously. Instead of a dead-air pause while it deliberates, the model uses early verbal cues — “Let me check that…” — to acknowledge a prompt naturally, then narrates its progress live as it works through multi-step background tasks. In Google’s demos it converts raw sketches into functional React components with near real-time voice feedback, and coordinates multi-step bookings with asynchronous function calls without ever breaking the conversational thread.

The numbers

Extended Thinking arrives with a genuinely strong benchmark sheet. Google says it captures the #1 overall spot on Artificial Analysis’ Speech-to-Speech Quality Index with a score of 82.6, leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark, and scores 97.7% on Big Bench Audio for audio-grounded reasoning — while, in Google’s words, “maintaining a highly competitive price point compared to other frontier models.”

The base 3.8 Live model takes second place in the Speech Agent Arena, and Google says both models push the Pareto frontier on ServiceNow’s EVA-Bench — a benchmark for evaluating voice agents on complex enterprise workflows — successfully balancing accuracy with conversational quality. That benchmark runs on the Live API on the Gemini Enterprise Agent Platform, which signals who Google sees as the real customer here: enterprises deploying production voice agents, not just consumers chatting with the Gemini app.

A polyglot that keeps up

Two quieter details deserve attention. First, Gemini 3.8 Live automatically detects and transitions between 97 supported languages mid-conversation — no menu, no restart, just a seamless switch when the user does. Second, every audio output across Google’s AI products is watermarked with SynthID, an imperceptible watermark woven directly into the audio signal so AI-generated speech remains detectable — an increasingly non-negotiable feature as voice cloning incidents pile up. A model card documenting Google’s safety approach ships alongside the release.

Where you can use it today

The rollout is unusually broad for a day-one launch:

  • Developers: both models are live now in the Gemini API and Google AI Studio.
  • Enterprises: private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience “coming soon.”
  • Everyone: Extended Thinking is rolling into Gemini Live today, plus Docs Live, Gmail Live, and Keep Live for Google AI Pro and Ultra subscribers in Workspace; the base 3.8 Live model powers Search Live.

Google is also leaning on the ecosystem to carry these models into production. Developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents now handle the real-time media streaming plumbing on top of the Gemini Live API, while early partners Salesforce, Genspark, and Lumeris are already building on the new models, citing latency, fluidity, and tool-calling as the differentiators.

Why this matters

The industry’s center of gravity has been shifting from chatbots to agents — models that act, not just answer — and the voice-agent corner of that shift has been the hardest to crack. Running a vision-capable, tool-calling agent over a live audio stream imposes brutal latency budgets: every reasoning step competes with the natural rhythm of human conversation. The pipelined architectures most products use can’t hide that. A model that genuinely deliberates while speaking, narrating its own progress instead of going silent, is a qualitatively different product — closer to a competent colleague on a call than a form with a microphone button.

It also sharpens the competitive picture. OpenAI has pushed its realtime voice stack hard, and Amazon, Microsoft, and a wave of startups are all chasing production voice agents. Google’s claim to the top of Artificial Analysis’ speech-to-speech index — if it holds up under independent testing — plus aggressive pricing and same-day availability across consumer surfaces, Workspace, and the API is a statement that it intends to own this layer the way it once owned search. Following March’s 3.1 Flash Live and last month’s Gemini 3.8 Flash and Flash Cyber launches, the cadence makes clear that live audio is no longer a side experiment at DeepMind; it’s a primary front in the model race.

The open questions are the usual ones: how the benchmarks translate to messy real-world deployments, whether “competitive pricing” survives contact with enterprise volume, and whether users actually want assistants that think out loud. But the bottleneck that made voice agents feel dumb — the inability to reason without going mute — just got meaningfully narrower.

Coverage based on Google’s official announcement and reporting from 9to5Google and Unite.ai; benchmarks as reported by Google.