Creative AI / The AI briefing

Google DeepMind launches Gemini 3.8 text-to-speech models for custom voice generation

Explore Gemini 3.8's advanced tools for creating lifelike voices tailored for games, audiobooks, and more.

Image accompanying the original report at Google DeepMind
From Google DeepMind. Product artwork from the source publication.

Google DeepMind has launched Gemini 3.8 Flash TTS and Flash-Lite TTS, allowing creators to generate expressive, customizable voices for applications like audiobooks and games.

What was announced

Google DeepMind unveiled two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS. These models allow users to generate fully customizable voices, offering options to modify accents and emotional tone. This level of personalization aims to enrich audio experiences in various applications such as audiobooks, gaming, and podcasts.

Developers and content creators can utilize these models through platforms like Google AI Studio and Google Vids, facilitating a dynamic way to produce audio that sounds natural and expressive. The tools are designed to enhance user engagement, making it easier to create unique character voices and improve scene dialogue.

Limits and availability

While Gemini 3.8 offers a significant upgrade in text-to-speech technology, potential limitations include the challenge of ensuring these voices are used ethically. Google's implementation of safety features aims to address misuse, but the effectiveness of these safeguards remains to be seen.

The models are part of Google DeepMind's expanding Gemini Audio family and are currently accessible via the Gemini API and enterprise solutions, allowing businesses to leverage them for customized audio solutions.

Original source

This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.

Read the original at Google DeepMind

← Back to all AI news