Thirty seconds.
That is how much audio Google's new text-to-speech model needs to "recreate consistent vocal profiles." Half a minute. A voice note. A clip from a school play. The outgoing message on a voicemail.
Google announced Gemini 3.8 Flash TTS and a cheaper Flash-Lite version this week, rolling out to developers through the Gemini API and Google AI Studio, with enterprise access "coming soon."

What It Does
Voice replication. "Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use."
Voice invention. Build a character from a written description - role, accent, vocal characteristics - across more than 100 languages and dialects. Google's own examples include a high-energy DJ from Melbourne and a Japanese dragon. The library goes from 30 preset voices to 2,000+, including regional varieties like Mexican Spanish, Quebec French and Scots English.
Direction. This is the part that makes it feel different from older text-to-speech. You direct the performance line by line: pacing, emotion, dialect shifts, and non-verbal cues - <laughs>, <sigh>, <gasp> - plus active-listening interjections like |mhm| and |yeah|, for what Google calls "precise comedic timing and reaction beats."
It also does two-speaker scenes from a single script, and holds a voice steady across hours of audio with "minimal speaker drift."
If you have ever noticed that synthetic narration never quite breathes, this is the release aimed squarely at that.
The Safeguard Most Coverage Will Skip
Here is where the honest version of this story diverges from the one you'll see elsewhere.
Google gates replication. From the announcement:
"For voice replication our system leverages consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created."
And every clip is marked:
"Every audio clip generated by our Gemini Audio models is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated speech remains detectable to help prevent misinformation."
Plus C2PA credentials for provenance.
So, as described: you cannot upload thirty seconds of a stranger's TikTok and receive their voice. A consent recording from the actual owner, matching the reference sample, is required first.
We are stating that plainly because the alternative version - "Google ships impersonation machine" - is going to be everywhere, and it is not what the page says. Consent verification plus watermarking plus provenance credentials is more than a lot of voice products ship with, and Google deserves the credit.
And Now the Limit
What follows is our analysis, not Google's claim.
Watermarking and provenance credentials are built for a particular shape of problem: media that somebody can inspect. A video uploaded to a platform. A clip a newsroom can run through a detector. A file with time to sit still while software examines it.
They work because there is a downstream moment with tooling in it.
The most common harm from synthetic voice does not have that moment.
It is a phone call. It is real-time. It is not recorded. It is decided in about fifteen seconds by a frightened person who has been told their child has been in an accident, or their grandchild is in a police station, and that money has to move now.
No watermark is legible in that channel. There is no upload, no scan, no detector, no second opinion. SynthID is doing real work for the information ecosystem, and approximately none for your mother at 11pm.
There is a second limit worth naming. Google's consent gate governs Google's product. It does not govern the wider field, where broadly comparable capability exists in tools with no equivalent check. The capability described in this announcement is where the whole industry is heading. The safeguards are one company's choice about their own product.
And one open question we cannot answer: the consent gate is a spoken recording that must match the reference speaker. Whether a synthesised consent recording would pass it is not addressed on the page, and nobody outside Google has tested it. We are raising that as a question, not an assertion.
What This Is Actually For
It would be dishonest to write this up as a fraud tool. The stated uses are substantial and mostly unglamorous.
Dubbing and localisation - partners include Linguana and Ollang, aimed at getting content into languages it currently never reaches. Audiobooks and podcasts - Wondercraft. Games and interactive media, where a small studio can now voice a cast it could never have afforded. Accessibility, for people who have lost their voice or never had the use of one. Voice agents, which is where most of the commercial money is.
The same property makes all of it work and makes the risk real: the output is no longer identifiably machine-made by ear. That was always going to arrive. It has.
The One Thing To Do This Week
Not "audit your digital footprint." Your child's voice is already online, in a school video, a group chat, a gaming session, a voicemail. Thirty seconds is not a high bar and you are not going to get under it.
Instead:
Agree a family safe word.
One word or short phrase, known only inside the household, never texted or posted. Anyone can ask for it on any call involving money, travel, an emergency, or a request to keep something secret. If the caller can't produce it, you hang up and ring the person back on the number you already have.
Some specifics that make it actually work:
Tell the older generation first. Grandparents are the most-targeted group for voice-based fraud and the least likely to have heard this advice.
Don't choose something guessable. Not a pet, not a street, not a birthday - all findable.
Practise it once. The failure mode isn't forgetting the word; it's feeling too silly to ask for it while someone is crying down the phone. One rehearsal removes that.
Give kids permission to hang up. A child who has been told to always be polite to adults will not end a call on someone using their parent's voice. Say explicitly that hanging up and calling back is correct behaviour.
Agree a second channel. "I'll text you and you reply" defeats almost all of this, because the attacker has one voice and not your accounts.
This costs a five-minute conversation at dinner and it is the only control in this whole story that works at the speed the attack works.
The Part Worth Remembering
Google built careful safeguards, and they are aimed at the problem a model provider can actually solve: keeping synthetic audio detectable in the record.
The problem they cannot solve is that recognising someone's voice has stopped being evidence that it is them. That was a piece of social infrastructure we all used, every day, without noticing we were relying on it.
It is gone. Replacing it costs one word, agreed in advance.
Source: Google, "Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS," blog.google. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/
This is a vendor product announcement, not independent reporting, and none of its capability or safety claims have been externally verified. Quotations are verbatim. The page carries no visible publish date; secondary summaries report 23 September 2026. Benchmark placements are third-party benchmarks selected and reported by Google. The analysis of watermarking versus live-channel fraud, the open question about the consent gate, and the family guidance are ours.
Disclosure: the AI by Age video narration is itself AI-generated. We are reporting on synthetic voice using synthetic voice, and think you should know that.
Want practical AI guidance for parents and educators every week? Subscribe: https://www.aibyage.com/?modal=signup&utm_source=beehiiv&utm_medium=newsletter