HeadlinesBriefing favicon HeadlinesBriefing.com

Google Debuts Gemini 3.8 Expressive Text-to-Speech Models

Google DeepMind Blog •
×

Google DeepMind has introduced two new text-to-speech models to the Gemini family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models transform voice generation from static presets into a dynamic creative studio, enabling creators, developers, and enterprises to craft richer, more expressive audio experiences. Gemini 3.8 Flash TTS is built for deep creative direction and character design, allowing users to generate custom voices from scratch via natural language prompts across gaming, audiobooks, podcasts, and interactive media. Gemini 3.8 Flash-Lite TTS focuses on high-volume, cost-efficient scale, optimized for dubbing, voice agents, and expressive content with fine-grained tone and pacing control.

The models expand the Gemini Audio family, following releases like 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. Users can scale from 30 original voices to an infinite library, accessing over 2,000 production-ready voices with broad language coverage including Mexican Spanish, Quebec French, and Scots English. Voice replication is supported from just a 30-second sample, backed by consent verification, Synth ID watermarking, and C2PA credentials.

Both models offer precise line-by-line performance control, long-form generation with minimal speaker drift, native two-speaker scene staging, and scripted vocal bursts for realistic conversational texture. Voice remixing capabilities are also coming soon, allowing timbre, pitch, pace, and accent fine-tuning via prompts.