What Google announced
Google introduced two new text-to-speech models on September 23, 2026: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The announcement, published on Google's official blog by Leland Rechis, Group Product Manager, and Alan Cowen, Director of Research Science, on behalf of the Gemini Audio Team, describes the models as moving voice generation "from static presets into a dynamic creative studio."
For anyone who produces audio, that framing matters. Until now, most synthetic speech has meant choosing from a fixed list of preset voices. Google's new pitch is different: design a voice from a text description, replicate an existing voice from a short recording, and then direct the performance line by line, the way a director coaches an actor.
Two models, two jobs
According to the announcement, Gemini 3.8 Flash TTS is built for what Google calls deep creative direction and character design. Users can create entirely new voices from scratch using natural language prompts, specifying role, accent, and voice characteristics across what Google says is more than 100 languages and dialects. The company gives examples ranging from a dramatic, fire-breathing dragon to a charismatic narrator with a distinct regional cadence.
Gemini 3.8 Flash-Lite TTS is positioned as the high-volume, cost-efficient option, aimed at large-scale dubbing, audio content creation, and expressive voice agents. The model card, published September 15, 2026, confirms the technical basics: both models take text input of up to 8K tokens and output audio with a 64K token output window, and both are based on Gemini 3 Pro.
Designed, not just selected
Several claimed capabilities stand out for non-specialists. Google says the models offer a library of more than 2,000 production-ready voices with regional varieties such as Mexican Spanish, Quebec French, and Scots English. Long-form generation is said to maintain voice quality and pacing across hours of continuous audio with minimal speaker drift, which is the property that matters most for audiobooks and podcasts.
The models also support native two-speaker scene staging from a single script, and scripted non-verbal cues such as <laughs>, <sigh>, <gasp> and active-listening interjections like |mhm| or |yeah|. These tags let a writer control conversational texture the way a screenplay controls action.
These are vendor descriptions of what the product is designed to do. They are not independent measurements of output quality, and this article does not treat them as such.
The 30-second question
The most consequential feature is voice replication. Google's exact wording on the announcement page is worth quoting in full:
Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.
Attribution: Google, Leland Rechis, Group Product Manager, and Alan Cowen, Director of Research Science, on behalf of the Gemini Audio Team, "Gemini 3.8 text-to-speech says hello," Google blog, September 23, 2026.
Safeguards and their limits
The safeguards Google describes work in layers. For voice replication, the announcement states that users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. SynthID, Google DeepMind's watermarking technology, embeds an inaudible watermark into every audio clip generated by the Gemini Audio models, per the announcement; the SynthID page describes audio watermarks generally as imperceptible to humans but detectable by SynthID's technology, and designed to survive common modifications such as compression or speed changes. C2PA credentials add cryptographic provenance metadata to the output.
These mechanisms have real limits, which Google does not dispute. Consent verification depends on a matching consent recording, which addresses impersonation through the tool itself but cannot reach misuse of cloned audio once it circulates. Watermarks make generated speech detectable, not impossible: detection requires the watermark to survive whatever transformations the audio undergoes, and the watermark identifies Google's models specifically, not AI audio generally. Voice replication in AI Studio is a cautious, rights-gated capability, but the underlying technical capacity, reproducing a voice from 30 seconds of audio, is now a shipped product rather than a research demo.
Where it is available, and where it is not
Availability is live for developers and consumers now, and pending for enterprises. Per the announcement, Gemini 3.8 Flash TTS is rolling out in the Gemini API and Google AI Studio for developers, in Gemini Notebook for general users, and is "coming soon" via API in Gemini Enterprise. Flash-Lite TTS follows the same pattern, with Google Vids as its consumer surface instead of Gemini Notebook.
One detail in a footnote deserves attention from the access angle: voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India. These jurisdictions have been among the most active in regulating biometric and voice data, and the carve-out means that the flagship feature of the release is geographically restricted at launch. Developers in excluded regions can still use the voice design and library features, but not replication of a specific person's voice through AI Studio.
Google also lists partner integrations with developer platforms including Agora, LiveKit, Pipecat, and Vercel, and with companies including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang. These are vendor-listed integrations; the scope and depth of each partner's implementation is not independently verified here.
Benchmarks: vendor claims, not verified facts
Google claims that Gemini 3.8 Flash TTS secured the number one overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and led in accent modeling at 60.8. The company also says both models took the top two spots on Hume AI's Overall Quality Index, and that blind human preference evaluations on Voice Arena placed them in top positions in several languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.
These are Google's claims, and this article could not independently verify the Hume leaderboard itself. The Hume AI Voice Design Benchmark page was not retrievable for verification at the time of writing, so readers should treat the rankings as self-reported vendor results attributed to a third-party benchmark. The model card confirms that Google evaluated the models against a range of benchmarks but does not reproduce the specific scores.
What the model card does say independently, within Google's own documentation, is a frontier safety assessment: Google states that the 3.8 Audio models do not have meaningful new capabilities or material performance increases compared to Gemini 3.7 Flash, and that they are not likely to reach any Tracked or Critical Capability Levels under Google's Frontier Safety Framework. Google also acknowledges general limitations shared with foundation models, including possible hallucinations and occasional slowness or timeouts.
What this means for people who work with sound
The practical consequence lands on people who already work with their voices. A small publisher that could never afford professional narration for an audiobook in six languages now has a plausible path to producing it. A dubbing studio can prototype accent and tone choices before booking studio time. A podcaster can fix a misread line without a reshoot.
The same capability rewrites the economics of voice work. If a narrator's own voice can be replicated from a 30-second sample, narrators can license their voices at scale, and so can anyone who obtains, or forges, the required consent recording. The question of who consents, and how consent is verified across jurisdictions with different voice-likeness laws, is no longer hypothetical. Google's own footnote, restricting replication in several major markets, is an implicit acknowledgment that the legal landscape is unsettled.
None of this means the generated voices are indistinguishable from human speech, a claim this article does not make. It means the direction of travel is clear: voice is becoming a designed, directed, replicable asset, with watermarking and consent checks as the current, imperfect guardrails.
How we know what we know
The facts in this article come from Google's announcement, Google's model card, and Google's SynthID documentation, all retrieved and checked. The benchmark rankings, partner integrations, and quality characterizations are Google's own statements and, unless noted, are not independently verified. The Hume AI leaderboard page was not retrievable at the time of writing. The interpretation of consent, watermarking limits, and industry impact is analysis by the author.
