OFICIAL Google DeepMind Blog AI & Software · Jun 09, 2026

Fluid, natural voice translation with Gemini 3.5 Live Translate

In brief · 4 sentences
Based on Google DeepMind Blog · Jun 09, 2026

Google introduced Gemini 3.5 Live Translate, an audio model enabling near real-time speech-to-speech translation in over 70 languages with natural intonation and minimal delay.

Fluid, natural voice translation with Gemini 3.5 Live Translate
Google DeepMind Blog — Google
Key points
·
Main topic: fluid, natural voice translation with Gemini 3.5 Live Translate.
·
Category affected: AI and software.
·
Figures mentioned: 3.5, 70, 10 million.
·
The information comes from an official source.
·
The next step is to watch availability, pricing and real-world impact.

The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.

Google has launched Gemini 3.5 Live Translate, a new audio model designed for live speech-to-speech translation across more than 70 languages. The system processes speech as it is streamed, generating translated audio continuously rather than waiting for pauses. This approach aims to reduce awkward silences while maintaining accuracy, with output typically lagging just a few seconds behind the speaker. The model automatically detects languages and adapts to multilingual inputs without manual configuration, making it suitable for real-time applications like calls and meetings.

The technology is now available in a private preview for select Google Workspace business customers, with a broader rollout planned later in the year. It will also be integrated into the Google Translate app for Android and iOS users globally. The Live Translate feature supports headphone-based translation, preserving the speaker’s tone and pacing. Android users will additionally gain a new ‘listening mode,’ allowing translations to stream directly through the phone’s earpiece for discreet use.

Developers can access the Gemini Live API to build voice translation applications, with support from platforms like Agora, Fishjam, LiveKit, Pipecat, and Vision Agents. These integrations handle real-time media streaming, enabling developers to focus on user experience. Early partners, including Grab, are testing the model to facilitate multilingual communication between drivers and travelers, handling over 10 million voice calls monthly.

All audio generated by the model is watermarked with SynthID, an imperceptible watermark embedded in the output to help detect AI-generated content and mitigate misinformation. Google emphasizes its commitment to safety and responsibility, providing a model card for further details on its approach. The feature aligns with Google’s ongoing efforts to expand translation services, building on two decades of machine learning advancements.

Original source → Deals on Clipraptor.com →
Extracted signals · detected in the story
FluidGeminiLive Translate. GeminiLive TranslateGoogle AI StudioGoogle TranslateGoogle Meet.. GeminiYourTwentyToday3.57010 million