Gemini 3.8 text-to-speech says hello
Google introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, enabling highly customizable and expressive voice generation for creators and developers across multiple platforms.
Google unveiled two advanced text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, designed to produce more expressive and natural-sounding audio. These models allow users to generate custom character voices and direct scene dialogue directly within tools like Google AI Studio, the Gemini API, and Google Vids. Creators can now fine-tune accents, emotional tones, and vocal styles, making the technology suitable for audiobooks, games, and podcasts. The models also include safety features to promote responsible use of generated voices.
The new models expand the existing Gemini Audio family, complementing tools like 3.5 Live Translate and 3.8 Live. They enable an infinite library of voices, from original characters to brand ambassadors, with precise control over delivery. Developers can integrate these models into platforms such as Agora, LiveKit, and Vercel to build custom audio experiences. The models support over 100 languages, facilitating multilingual voice creation for global audiences.
Gemini 3.8 Flash TTS leads in voice customization, securing the top spot on Hume AI’s Voice Design Benchmark with a score of 71.4. It also ranks first in accent modeling at 60.8 and holds the #1 and #2 positions on Hume AI’s Overall Quality Index for both Flash TTS and Flash-Lite TTS. Improvements include better performance in long-form content and dual-speaker screenplay control compared to the previous 3.1 Flash TTS model.
Google implemented strict safeguards for voice replication, requiring verbal consent recordings from voice owners before creating a replica. All generated audio is watermarked with SynthID to ensure transparency and help prevent misinformation. Developers can access these capabilities in Google AI Studio, while partners like Figma, HeyGen, and Wondercraft integrate the models for dubbing, localization, and conversational agents.