← All posts / Tools

Google Puts a Face on Gemini: Live Avatar Ships to Enterprises in 97 Languages

Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise — real-time lip-synced video avatars, 97-language switching, background tool calls, and SynthID watermarks on every generated frame.

Google Puts a Face on Gemini: Live Avatar Ships to Enterprises in 97 Languages

Ten days after Google unveiled Gemini 3.8 Live — its speech-native dialogue model that reasons mid-sentence — the company has completed the picture, literally. On September 25, 2026, Gemini 3.8 Live with Live Avatar became generally available in Gemini Enterprise, adding a real-time animated face to the voice. An AI that not only answers but looks back at you, lip-syncing its reply and switching languages mid-conversation, is now an enterprise production feature rather than a conference demo.

First previewed at Google Cloud Next 2026, the technology pairs the native speech-to-speech foundation of Gemini 3.8 Live with near-real-time video generation. The result is a conversational video agent that runs across web browsers, mobile apps, and interactive kiosks — the kiosk form factor being an obvious hint at where Google expects this to land first: retail counters, hospital check-in desks, bank branches, and airport gates.

What Live Avatar actually does

The feature set reads like a checklist of everything voice-only agents have historically gotten wrong:

Video avatars with synchronized lip-syncing. The avatar’s mouth, and per early hands-on reports its facial expressions, move in sync with the generated speech. When the conversation switches languages, the lip-syncing adapts on the fly — no re-render, no frozen frame while the model switches gears.

97 languages with automatic detection. Gemini 3.8 Live understands and speaks 97 languages, detecting the caller’s language automatically. For global support operations, that collapses what used to be a routing nightmare across regional voice models into a single deployment.

Background tool calling. The model executes tools and API calls mid-conversation without going silent. It acknowledges a request, keeps chatting naturally while the task runs in the background, and reports back when it’s done — the difference between an assistant that puts you on hold and one that multitasks.

Live visual understanding. The model processes live camera feeds and screen shares alongside audio, simultaneously. Show it a dented fender on camera while describing the accident, and it sees and hears both streams at once.

Interruption recovery that holds context. Because the dialogue is native speech-to-speech rather than a chained pipeline, interruptions don’t drop conversation state or in-flight backend transactions — a chronic failure mode of earlier voice agents.

Trust rails: allowlists and watermarks

The more interesting details are the guardrails. Customers can deploy from a library of curated, pre-built avatars out of the box, but custom avatar creation — building a face from a single reference photo and audio sample — is gated behind a strict enterprise allowlisting and verification process. You cannot, as an individual developer, spin up a photorealistic replica of anyone today.

And every generated audio and video stream carries an imperceptible SynthID watermark, so content produced by a Live Avatar remains machine-verifiable as AI-generated even after it leaves Google’s servers. In a year when synthetic video has become a political and legal flashpoint, shipping a talking human face without provenance marking would have been indefensible; Google clearly judged the same.

The enterprise posture extends to infrastructure: US and EU endpoints, provisioned throughput options, and the data-governance commitments that regulated industries demand before letting a video agent touch customer conversations. Gemini 3.8 Live Extended Thinking — the slower, deeper-reasoning sibling — remains in private preview.

Customers already running it

Google’s launch post came with named deployments, not just partner logos. Cox Automotive built a shopping assistant for Autotrader that uses live screen-highlighting and tool-calling to guide car shoppers through search, comparison, and financing in natural conversation. “Shoppers increasingly expect to describe what they need in their own words rather than work through filters and menus,” said Marianne Johnson, EVP and Chief Product Officer at Cox Automotive.

Equal AI, which handles over a million live calls daily across nine Indian languages, credits the new version with improved interruption handling, multilingual conversations, and tool-call reliability. “This AI doesn’t just answer calls; it gets things done for you,” said CEO Akhilesh Dhar. Salesforce is pairing Gemini 3.8 Live with Agentforce for customer service, and engineers at Specs highlighted the improved voice-activity detection and latency as concrete wins over earlier Live generations.

The demos tell the same story from the builder’s side. One shows an insurance claims intake agent: you talk and show the damage on camera while a claim notebook fills itself in, and an agent team built with Google’s Agent Development Kit (ADK) checks the policy, applies intake rules, and assembles the adjuster packet in the background — with the open-source code published for anyone to replicate.

Why a face changes the economics

Voice AI has been trapped in a commodity metrics race — latency milliseconds and cost per minute. Google’s product positioning explicitly says its priority has shifted to interaction quality, and the avatar is the sharpest expression of that bet. A face does two things a voice cannot.

First, it moves AI interfaces to places screens already exist. Every kiosk, checkout terminal, and waiting-room display becomes a potential agent deployment. Second, and more subtly, visual presence changes user behavior: people interrupt less rudely, listen longer, and complete more transactions when an interface presents social cues. Whether that effect survives contact with real users at scale is an open empirical question — but it’s the question enterprises will now be testing, because the feature is live and the API is documented.

There are legitimate concerns in the other column. The uncanny valley remains uncanny for a nontrivial share of users. Photorealistic enterprise avatars raise impersonation risks that an allowlist mitigates but does not eliminate. And SynthID verification depends on downstream parties actually checking — a watermark nobody scans protects nobody.

The competitive read

The move puts pressure directly on the emerging “digital human” market — startups selling avatars as a layer bolted on top of third-party voice models — because Google is now bundling the face, the voice, the reasoning, the tool-calling, and the watermarking into one governed API. It also draws a contrast with rivals: OpenAI’s voice efforts remain audio-first, and Anthropic has stayed out of real-time consumer modalities entirely. Google is betting that the next interface battleground is a face, and it has shipped first at enterprise grade.

For developers, the pieces are public today: the model in Gemini Enterprise, Live API documentation on Google’s docs site, a Skill for real-time bidirectional streaming apps, and modular audio alternatives (Gemini 3.5 Transcribe, 3.5 Live Translate, and the 3.8 Flash TTS family) for cases that don’t need the full avatar. Custom avatar access and provisioned throughput route through Google Cloud sales.

Ten days separated the voice from the face. The pace is the message: speech-native models went from announcement to production avatar platform in under two weeks, and every enterprise contact surface with a screen is now in play.


Sources