Create your own voices with Gemini 3.8 text-to-speech

Categories: AI, Product

Summary

Google's Gemini 3.8 text-to-speech now offers 1,000+ ready-made voices plus custom voice creation with verbal identity verification, enabling builders to generate expressive, consistent audio output on demand without licensing friction.

Key Takeaways

  1. Access to 1,000+ pre-built voices eliminates the need to build custom voice libraries, accelerating time-to-market for audio products and reducing development complexity for founders scaling voice features.
  2. Voice customization in real-time (pitch, tone depth adjustments) enables A/B testing different character voices without re-recording, reducing iteration cycles for creative and narrative content production.
  3. Custom voice cloning with 'verbal identity check' provides a safety mechanism for personal voice recreation, addressing privacy concerns while enabling personalized audio experiences without explicit consent complexities.
  4. Multi-speaker scene direction allows non-technical creators to assign and manage different voices to dialogue characters, lowering the barrier for indie developers and small studios to produce professional audiovisual content.
  5. Save and reuse favorite voices across projects creates a voice asset library effect, reducing cognitive load and ensuring consistent brand voice across multiple video/audio outputs for content creators.

Related topics

Transcript Excerpt

I need to make a promo video for our new text to speech model. Let's start with a voice over. >> Introducing text to speech. >> No, more like this. >> Introducing text to speech. Transform simple prompts into consistent expressive voice output on demand. >> Okay, now let's add in another voice for this [music] dialogue scene in the script. >> Or choose from over a thousand ready to go voices. Okay, but what if we had that voice a little bit deeper >> and change it however you want? >> Okay, perfect. Let's drop these voices into the dialogue scene. >> You can pick and drop your favorite save voices. >> And you can direct them. >> Uh-huh. In a multi- speakeraker scene. >> Actually, maybe it should be my voice. The lighthouse keeper watched the silver horizon as the morning fog slowly lifted …

More from Google DeepMind

Featured in