OFICIAL Google DeepMind Blog

Intelligent transcription with Gemini 3.5 Transcribe

What happened
Based on Google DeepMind Blog · Aug 26, 2026

Google introduced Gemini 3.5 Transcribe, a new speech-to-text model designed for precise real-time transcription with reduced latency and improved accuracy across multiple languages.

Intelligent transcription with Gemini 3.5 Transcribe
Google DeepMind Blog — Google
Key points
·
Google its latest speech-to-text model designed for precise and intelligent real-time transcription.
·
Today, the company is introducing Gemini 3.5 Transcribe, its most precise speech-to-text model yet, designed for intelligent voice interactions.
·
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
·
Across its products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS.
Key numbers
·
Performance benchmarks show a 70% reduction in time to final transcription and lower word error rates in both streaming and non-streaming modes.
·
50% in streaming mode and 5.
·
04% in non-streaming use cases across top languages and locales.

Google has launched Gemini 3.5 Transcribe, a speech-to-text model aimed at improving transcription accuracy and reducing latency compared to its predecessor, Chirp 3. The model is now available in the Gemini API through Google AI Studio and the Gemini Enterprise Agent Platform, enabling developers to integrate advanced voice capabilities into applications. It supports real-time captioning, voice agents, and post-call analytics, with features like multi-speaker attribution and word-level timestamps. Performance benchmarks show a 70% reduction in time to final transcription and lower word error rates in both streaming and non-streaming modes.

The model enhances transcription quality by handling background noise, complex jargon, and speech disfluencies, converting raw audio into polished text. It is already integrated into consumer products like the Gemini app and Android, where features such as Rambler on Android and voice search in the Gemini app on macOS demonstrate its practical applications. Developers can leverage these capabilities to build voice-driven interfaces with improved natural language understanding and custom vocabulary recognition.

Gemini 3.5 Transcribe introduces context-aware features across Google services, including Gboard, Antigravity, and Chrome, to make interactions more intuitive. It enables inline edits, intent recognition, and screen context integration, allowing users to perform tasks like file analysis, image generation, and searches using voice commands. The model’s multilingual performance, as measured by the FLEURS benchmark, shows a word error rate of 5.50% in streaming mode and 5.04% in non-streaming use cases across top languages and locales.

The model is supported by platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents, which facilitate the development of high-performance voice interfaces. Companies like Vivo, Intellitek Health, and Lingopal have reported positive feedback on its latency, accuracy, and language support. Google highlights its use in consumer and enterprise applications, emphasizing improved workflow integration and user experience through advanced transcription capabilities.

Original source → Deals on Clipraptor.com →