OFICIAL The Keyword

Intelligent transcription with Gemini 3.5 Transcribe

What happened
Based on The Keyword · Aug 26, 2026

Google unveiled Gemini 3.5 Transcribe, a new speech-to-text model offering improved accuracy, multilingual support, and reduced latency for real-time transcription and voice-driven applications.

Intelligent transcription with Gemini 3.5 Transcribe
The Keyword — Google
Key points
·
Google its latest speech-to-text model designed for precise and intelligent real-time transcription.
·
Today, the company is introducing Gemini 3.5 Transcribe, its most precise speech-to-text model yet, designed for intelligent voice interactions.
·
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
·
Across its products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS.
Key numbers
·
According to Artificial Analysis, it reduces transcription time by 70% compared to its predecessor.
·
50% word error rate in streaming mode and 5.
·
04% in non-streaming scenarios across multiple languages and locales.

Google introduced Gemini 3.5 Transcribe, a speech-to-text model designed to handle background noise, jargon, and disfluencies more effectively than prior versions. The model converts raw audio into polished, formatted text and is already integrated into consumer products like the Gemini app and Android’s Rambler feature. Developers can now access it via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform to build voice agents, captioning tools, or analytics pipelines.

Gemini 3.5 Transcribe improves transcription accuracy with multi-speaker attribution and word-level timestamps, addressing gaps in earlier models like Chirp 3. According to Artificial Analysis, it reduces transcription time by 70% compared to its predecessor. On the FLEURS benchmark, it achieves a 5.50% word error rate in streaming mode and 5.04% in non-streaming scenarios across multiple languages and locales.

The model enhances user experience across Google’s ecosystem, including Gboard, Antigravity, and Chrome, by capturing nuances and intent through context-aware understanding. In the Gemini app on macOS, users can analyze files, generate images, and search using voice commands. On Android, Rambler automatically removes filler words and cleans up speech, while Antigravity leverages screen context for improved transcription accuracy.

Developers can integrate Gemini 3.5 Transcribe via the Gemini Live API to build voice-driven interfaces, with support from platforms like Agora, Fishjam, and LangChain. Companies such as Vivo, Intellitek Health, and Lingopal have reported positive feedback on its latency, accuracy, and language support. The model is now available for broader deployment in developer workflows.

Original source → Deals on Clipraptor.com →