Create your own voices with Gemini 3.8 text-to-speech
Summary
Google's Gemini 3.8 text-to-speech now offers 1,000+ ready-made voices plus custom voice creation with verbal identity verification, enabling builders to generate expressive, consistent audio output on demand without licensing friction.
Key Takeaways
- Access to 1,000+ pre-built voices eliminates the need to build custom voice libraries, accelerating time-to-market for audio products and reducing development complexity for founders scaling voice features.
- Voice customization in real-time (pitch, tone depth adjustments) enables A/B testing different character voices without re-recording, reducing iteration cycles for creative and narrative content production.
- Custom voice cloning with 'verbal identity check' provides a safety mechanism for personal voice recreation, addressing privacy concerns while enabling personalized audio experiences without explicit consent complexities.
- Multi-speaker scene direction allows non-technical creators to assign and manage different voices to dialogue characters, lowering the barrier for indie developers and small studios to produce professional audiovisual content.
- Save and reuse favorite voices across projects creates a voice asset library effect, reducing cognitive load and ensuring consistent brand voice across multiple video/audio outputs for content creators.
Related topics
Transcript Excerpt
I need to make a promo video for our new text to speech model. Let's start with a voice over. >> Introducing text to speech. >> No, more like this. >> Introducing text to speech. Transform simple prompts into consistent expressive voice output on demand. >> Okay, now let's add in another voice for this [music] dialogue scene in the script. >> Or choose from over a thousand ready to go voices. Okay, but what if we had that voice a little bit deeper >> and change it however you want? >> Okay, perfect. Let's drop these voices into the dialogue scene. >> You can pick and drop your favorite save voices. >> And you can direct them. >> Uh-huh. In a multi- speakeraker scene. >> Actually, maybe it should be my voice. The lighthouse keeper watched the silver horizon as the morning fog slowly lifted …
More from Google DeepMind
- From deepfakes to DNA: the science of watermarking AI
- AlphaGenome Atlas: Understanding the human genome
- Robots working together with Gemini Robotics 2
- WeatherNext 3: More accurate, timely, and local weather forecasts
- The mathematics of AI uncertainty