MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world
Microsoft introduces MAI-Transcribe-2, a speech recognition model surpassing competitors in speed, accuracy, and cost efficiency, with new features like diarization and word-level timestamps.
Microsoft has unveiled MAI-Transcribe-2, a speech recognition model that outperforms leading alternatives such as Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2 in both accuracy and efficiency. The model introduces features like diarization, configurable transcription styles, and word-level timestamps, expanding its practical applications across sectors including healthcare, legal, accessibility, and captioning. It ranks first on the FLEURS benchmark across 60 languages with a 5.2% average Word-Error-Rate and defines the Pareto Frontier for accuracy and latency on Artificial Analysis.
MAI-Transcribe-2 demonstrates significant performance advantages, achieving 10 times the speed of OpenAI’s GPT-Transcribe, 7 times that of ElevenLabs’ Scribe v2, and 5 times that of Gemini 3.5 Transcribe while maintaining higher accuracy. Its efficiency is particularly beneficial for low-latency transcription needs, offering developers a clear performance edge over competitors in real-world deployments.
The model’s multilingual capabilities are highlighted by consistent high accuracy across all tested languages in the FLEURS benchmark, enabling developers to use a single model for transcription across multiple languages. This reduces complexity and potential GPU utilization issues, simplifying deployment for applications requiring broad language support.
MAI-Transcribe-2 is priced at $0.10 per hour during a limited-time launch offer, positioning it as the most cost-effective option in the market. The model is accessible for demonstration through Microsoft Foundry, MAI Playground, and Open Router, with availability expanding as part of Microsoft’s broader AI initiatives.