Google DeepMind has introduced two new text-to-speech (TTS) models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, designed for highly expressive audio generation. These models aim to transform voice generation from static presets into a dynamic creative tool, enabling the creation of custom character voices and direct scene dialogue.
PLUS ULTRAModel ReleasesGoogleGoogle DeepMindGemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS
Google DeepMind introduces Gemini 3.8 Flash TTS for expressive audio generation
PLUS ULTRA by Amenoyomi
The new models allow for granular control over vocal delivery, including accent and emotional tone. According to Google, these models offer significant improvements in use cases such as long-form content and dual-speaker screenplay control compared to the previous Gemini 3.1 Flash TTS. Gemini 3.8 Flash TTS achieved high scores on Hume AI’s Voice Design Benchmark and lead in accent modeling.
In blind human preference evaluations on Voice Arena, both models secured top positions in key global languages, including Japanese. The models support over 100 languages and are available through Google AI Studio and the Gemini API.
To ensure responsible use, Google has implemented safety features, including consent verification for voice replication. All audio clips generated by Gemini Audio models are watermarked with SynthID, an imperceptible watermark designed to make AI-generated speech detectable.
PLUS ULTRAby Amenoyomi
The transition from static presets to a prompt-based system changes how vocal identities are created. Instead of selecting from a limited set of 30 original voices, users can now use natural language prompts to generate bespoke vocal personas from scratch. This allows for the creation of an infinite library of voices, ranging from a high-energy DJ from Melbourne to a monotone robot or a Japanese dragon.
Beyond the initial voice design, the models provide granular control over the performance of each line. Creators can direct specific accents and emotional tones to ensure the delivery matches the intended scene. This includes the ability to manage natural turn-taking in dialogue, enabling the transformation of scripts into fully performed scenes through a dual-speaker screenplay editor.
These capabilities are supported by objective performance metrics in voice customization. Gemini 3.8 Flash TTS secured the top position on Hume AI’s Voice Design Benchmark with a score of 71.4 and led in accent modeling with a score of 60.8. Such precision in vocal identity and delivery is intended for professional creative workflows, including the production of audiobooks, games, and podcasts.
Sources
- Gemini 3.8 text-to-speech says hello (Google DeepMind Blog, 2026-09-23)
- Google Blog