Intelligent transcription with Gemini 3.5 Transcribe
Google unveiled Gemini 3.5 Transcribe, a new speech-to-text model offering improved accuracy, multilingual support, and reduced latency for real-time transcription and voice-driven applications.
Google introduced Gemini 3.5 Transcribe, a speech-to-text model designed to handle background noise, jargon, and disfluencies more effectively than prior versions. The model converts raw audio into polished, formatted text and is already integrated into consumer products like the Gemini app and Android’s Rambler feature. Developers can now access it via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform to build voice agents, captioning tools, or analytics pipelines.
Gemini 3.5 Transcribe improves transcription accuracy with multi-speaker attribution and word-level timestamps, addressing gaps in earlier models like Chirp 3. According to Artificial Analysis, it reduces transcription time by 70% compared to its predecessor. On the FLEURS benchmark, it achieves a 5.50% word error rate in streaming mode and 5.04% in non-streaming scenarios across multiple languages and locales.
The model enhances user experience across Google’s ecosystem, including Gboard, Antigravity, and Chrome, by capturing nuances and intent through context-aware understanding. In the Gemini app on macOS, users can analyze files, generate images, and search using voice commands. On Android, Rambler automatically removes filler words and cleans up speech, while Antigravity leverages screen context for improved transcription accuracy.
Developers can integrate Gemini 3.5 Transcribe via the Gemini Live API to build voice-driven interfaces, with support from platforms like Agora, Fishjam, and LangChain. Companies such as Vivo, Intellitek Health, and Lingopal have reported positive feedback on its latency, accuracy, and language support. The model is now available for broader deployment in developer workflows.