A Vocal Studio in an API: Google's Gemini 3.8 TTS Turns Text Prompts Into Directed Performances
Google's Gemini 3.8 Flash TTS and Flash-Lite TTS replace static voice presets with a promptable vocal studio — 2,000+ voices, 30-second voice replication with consent checks, line-by-line acting direction, and native two-speaker scenes, all watermarked with SynthID.
Voice generation has spent a decade stuck in the preset era: pick a voice from a dropdown, adjust speed and pitch, accept whatever expressiveness the engine deigns to deliver. On September 23, Google took a sledgehammer to that model. Gemini 3.8 Flash TTS and its cheaper sibling Gemini 3.8 Flash-Lite TTS, announced on the Google blog by Group Product Manager Leland Rechis and research director Alan Cowen, transform text-to-speech from a preset player into what Google describes as a “dynamic creative studio” — a vocal studio that lives inside an API.
The two models ship starting today in the Gemini API and Google AI Studio for developers, land in Gemini Notebook for consumers, and are “coming soon” to Gemini Enterprise, with Flash-Lite TTS also powering Google Vids. Together they complete a fast-growing Gemini Audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. Google is clearly assembling a full audio stack — and the new TTS pair is its most aggressive attack yet on the voice-generation market.
Two Models, Two Jobs
The split is deliberate. Gemini 3.8 Flash TTS is the flagship creative model, “engineered for studio-grade voice fidelity,” built for deep creative direction and character design. You use it to invent entirely new voices from natural-language descriptions — Google’s own demos include a high-energy DJ from Melbourne, a “super-tinny, monotone robot,” and a Japanese dragon — and then direct every performance line by line, with granular control over acting cues, pacing, dialect shifts, and backchanneling.
Gemini 3.8 Flash-Lite TTS is the workhorse: optimized for high-volume dubbing, bulk audio content, and expressive voice agents, where cost per minute matters more than bespoke character work. It keeps fine-grained control over tone and pacing but is priced for pipelines that generate hours of speech a day.
From 30 Voices to an Infinite Library
The most striking shift is scale. The previous generation offered 30 preset voices; the new family opens with a library of 2,000+ production-ready voices with broad language coverage, including regional varieties like Mexican Spanish, Quebec French, and Scots English. On top of that sits generative voice design: describe a role, an accent, and vocal characteristics in plain language — across more than 100 languages and dialects — and the model conjures a bespoke voice that doesn’t exist in any library.
Then there is voice replication. Provide a 30-second audio sample of your own voice, or a voice you have the rights to use, and the model recreates a consistent vocal profile from it. Google has wrapped this in the kind of paperwork the technology demands: a verbal consent recording from the voice owner must match the reference speaker before a voice can be created, every output carries a SynthID watermark woven imperceptibly into the audio, and clips ship with C2PA content credentials. Notably, voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland, or India — a quiet map of where the legal fight over voice rights currently stands. A “voice remixing” feature that lets you fine-tune timbre, pitch, pace, and accent of library voices is listed as coming soon.
Directing the Performance
Where the models genuinely break new ground is performance control. Both TTS variants accept stage directions inline in a script — from “a calm customer service agent” to “a whispered suspense scene” — and steer delivery accordingly. Long-form generation is built to hold voice quality, pacing, and character timbre across hours of continuous audio with minimal speaker drift, which is the failure mode that has historically made AI audiobooks unbearable.
Two features stand out for anyone who has tried to stage dialogue with older TTS systems. Native two-speaker scene staging directs multi-turn conversations from a single script while keeping both voices distinctly separated with natural turn-taking — no more stitching together separate generations and hoping the prosody lines up. And scripted vocal bursts let writers inject non-verbal texture using inline tags like <laughs>, <sigh>, and <gasp>, plus active-listening interjections like “mhm” and “yeah,” for precise comedic timing and reaction beats. That is acting direction, not speech synthesis.
The Numbers Behind the Voice
Google backed the launch with benchmark claims. Gemini 3.8 Flash TTS takes the #1 overall spot on Hume AI’s Voice Design Benchmark with a score of 71.4 and also leads in accent modeling at 60.8; the two models secure the #1 and #2 spots on Hume AI’s Overall Quality Index. In blind human preference evaluations on Voice Arena, they hold top positions in key global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi — with support for more than 100 languages overall. Google attributes major improvements in long-form content and dual-speaker screenplay control relative to Gemini 3.1 Flash TTS.
On price, the Gemini API lists Flash TTS at $0.50 per million text-input tokens and $9 per million audio-output tokens — widely worked out to roughly $0.81 per hour of generated speech, with introductory rates reported around 1.35 cents per minute for Flash and 0.9 cents for Flash-Lite through the end of the year. Flash-Lite undercuts that again for volume deployments. Third-party trackers have been quick to note what that does to incumbents: it puts direct pricing pressure on ElevenLabs and the boutique voice studios, especially for localization and dubbing work where the volume is enormous and the tolerance for per-minute premium pricing is not.
A Platform Play, Not Just a Model
The rollout details reveal the ambition. Developer platforms including Agora, LiveKit, Pipecat, and Vercel are wiring the models into real-time voice-agent stacks, and Google named Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang as launch partners integrating the TTS family for global dubbing, regional-accent localization, and conversational agents at scale. When OpenAI makes you call sales for a custom voice, as one industry commentator put it, Google is shipping a self-serve vocal studio with consent verification built in.
The strategic read is straightforward. Voice is becoming the interface for agents, and the quality bar for what passes as “natural” is rising fast. By combining generative voice design, rights-managed replication, watermarking, and two-speaker direction in a single API — and benchmarking aggressively to prove it — Google is positioning Gemini Audio as the default layer for the industry’s voice layer, the same way it chased the developer mindshare war in text models.
The preset era of text-to-speech is over. What replaces it looks less like a speech engine and more like a casting director.
Sources
- [1] https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/
- [2] https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts
- [3] https://aistudio.google.com/learn/gemini-3-8-flash-tts-developer-guide
- [4] https://www.marktechpost.com/2026/09/23/google-releases-gemini-3-8-flash-tts-and-flash-lite-tts-with-prompt-based-voice-design/
- [5] https://thenewstack.io/gemini-tts-voice-replication-api/