A two-model speech release

Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The Google announcement, dated 23 September, positions the first for detailed character design and the second for cost-efficient, high-volume speech generation. Both aim to give creators and developers more control over a performance than a simple choice from preset voices. The distinction matters for a studio making an audiobook and a service producing thousands of short voice responses each day.

Flash can create a voice from a natural-language description and direct individual lines using acting cues, pacing and dialect shifts. Google says it supports more than 100 languages and dialects and offers a library of over 2,000 production-ready voices. Flash-Lite is described as suitable for dubbing, audio content and voice agents where per-request economics matter.

These are text-to-speech systems, not a general licence to copy a person’s voice. The release makes voice replication possible from a short reference sample, but Google says it requires a matching verbal consent recording from the voice owner. A production team should verify rights and consent independently, even when a product supplies a technical check.

Directing a performance

Google describes line-by-line stage directions for calm or dramatic delivery, native two-speaker scenes and cues for laughter, sighs and conversational backchannels. For a podcast or interactive character, that can reduce the gap between a written script and a useful first recording. For customer support, the same controls could make an agent sound clearer, though tone should not conceal that the voice is synthetic.

Long-form output is another stated goal. The company says the models can maintain character and pacing across hours of audio with minimal drift. That is a vendor claim worth testing with an actual script, not a substitute for listening. A long audiobook may expose mispronunciations, inconsistent names or awkward transitions that a short demonstration never reaches.

A practical trial should include difficult proper names, Australian place names and mixed-language passages. Record whether the generated delivery follows the direction without changing the words, and how many edits are needed before publication. A voice that sounds expressive but alters a legal disclaimer or health instruction is unsuitable for that use.

Safeguards and evidence

Google says generated audio carries an imperceptible SynthID watermark and that voice replication uses consent verification. It also mentions C2PA credentials. These measures can support transparency, but they do not guarantee that every listener or platform can detect a synthetic clip, nor do they replace contracts with performers. Publishers should maintain a record of consent, source audio and intended use.

The company cites strong results on voice-design and quality evaluations, including Hume AI and Voice Arena. Comparative tests are signals, not a complete assessment of accessibility, naturalness in every language or how well a voice holds up at scale. Evaluation sets can differ from a broadcaster’s or developer’s scripts. Human listening remains essential.

Audio misuse is a real risk because a realistic voice can be mistaken for a person. A service using custom voices should consider disclosure to callers, restrictions on impersonation, and a process for reporting misuse. The stronger the voice-generation tool, the more important it is to make ownership and review traceable.

Where it is available now

Google says Flash TTS is rolling out in Gemini API and Google AI Studio and is available in Gemini Notebook, while Gemini Enterprise API access is coming soon. Flash-Lite is also beginning to roll out, with availability varying by surface. Voice remixing is explicitly labelled coming soon. Those distinctions should be preserved in planning; a feature shown in a demo is not necessarily deployable in every account.

For developers, the new audio playground in AI Studio provides a place to design voices and edit dual-speaker scripts before integrating through the API. Teams should confirm current quotas, latency, pricing and data handling at the actual endpoint they plan to use. A real-time voice agent will have different performance needs from an offline audiobook workflow.

The release is significant because it combines creative direction, broad language support and stated consent controls in a mainstream Gemini model family. Its practical value will be measured in consistent, rights-cleared audio that survives real scripts and production review, not just in the novelty of generating a convincing sample.

An editorial workflow should keep the original script beside each generated clip, identify the voice and model version, and retain the approval from the person whose voice is replicated. Check how the audio sounds through ordinary phone speakers and in noisy environments, not only through studio headphones. For multilingual material, have a fluent reviewer assess pronunciation and meaning. When a voice is updated, sample older episodes or messages to ensure the character remains recognisable. These steps cost time, but they prevent speed of generation from being confused with publishable quality, especially when the audience hears a familiar human voice.