OFICIAL Microsoft Source

New Microsoft AI models bring faster transcription and multilingual voices

What happened
Based on Microsoft Source · Oct 01, 2026

Microsoft introduced MAI-Transcribe-2-Streaming for real-time transcription in 60 languages and two new voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, optimized for multilingual conversational AI at low latency and cost.

New Microsoft AI models bring faster transcription and multilingual voices
Microsoft Source — Microsoft
Key points
·
MAI-Transcribe-2-Streaming provides real-time transcription in 60 languages with first hypotheses in just over 100ms
·
MAI-Voice-2.1-Flash generates 45s of audio with 150ms latency and costs $15 per 1M characters
·
MAI-Voice-2.1 supports 23 languages and 26 locales with a single voice maintaining native accents across languages
Key numbers
·
Microsoft unveiled MAI-Transcribe-2-Streaming, a real-time transcription service that generates low-latency transcripts in 60 languages with automatic language detection.
·
Internal tests show it delivers transcripts 2x faster than the nearest competitor for use cases like live dictation or subtitling.
·
1-Flash is designed for high-volume, latency-sensitive tasks and generates 45 seconds of audio with an end-to-end latency of 150 milliseconds.

Microsoft unveiled MAI-Transcribe-2-Streaming, a real-time transcription service that generates low-latency transcripts in 60 languages with automatic language detection. It ranks first for accuracy on Artificial Analysis for both final and partial transcripts while maintaining a balance between speed and precision. The service produces initial text hypotheses in just over 100 milliseconds, revising them as context accumulates, enabling applications to act on speech mid-sentence. Internal tests show it delivers transcripts 2x faster than the nearest competitor for use cases like live dictation or subtitling.

The new MAI-Voice-2.1 model supports 23 languages and 26 locales, allowing a single voice to speak multiple languages with native accents without changing speakers. This enables consistent branding across languages, such as a tutoring app switching languages mid-lesson or a multilingual assistant replying in the user’s language while retaining the same voice. The model includes built-in consent guardrails to prevent misuse and supports voice cloning with just a few seconds of reference audio.

MAI-Voice-2.1-Flash is designed for high-volume, latency-sensitive tasks and generates 45 seconds of audio with an end-to-end latency of 150 milliseconds. It offers 55% faster inference and is approximately 60% cheaper than comparable models, priced at $15 per 1 million characters. The model supports the same languages and cross-language speakers as MAI-Voice-2.1, making it suitable for pairing with MAI-Transcribe-2-Streaming in low-latency voice agent applications.

To demonstrate the models’ capabilities, Microsoft built Chatter, a live demo in the MAI Playground that showcases a voice agent using MAI-Transcribe-2-Streaming and MAI-Voice-2.1-Flash. The company highlighted its rapid compute expansion and ambitious roadmap, emphasizing its mission to deploy advanced AI models at scale with partners reaching billions of users.

Original source → Deals on Clipraptor.com →