Google released two text-to-speech models on September 23, 2026. Here's everything that matters in one place.
The two models
- Gemini 3.8 Flash TTS: built for creative direction and character voices.
- Gemini 3.8 Flash-Lite TTS: built for high-volume, low-cost generation.
What you can do
- Pick from 2,000+ production-ready voices in 100+ languages and dialects.
- Direct each line in plain English, and script vocal bursts like
<laughs>or<sigh>. - Generate long-form audio and two-speaker scenes.
Voice cloning rules
- A 30-second audio sample is enough.
- The voice owner must record spoken consent, and it has to match the sample's speaker.
- Not available in the UK, EEA, Switzerland, India, Illinois or Texas.
Safety
Every clip is watermarked with SynthID (inaudible) and ships with C2PA content credentials.
Where to try it
| Where | Flash TTS | Flash-Lite TTS |
|---|---|---|
| Gemini API / Google AI Studio | Today | Today |
| Gemini Enterprise | Coming soon | Coming soon |
| Consumer apps | Gemini Notebook | Google Vids |
Sources
blog.google
blog.google