Microsoft has released three new AI speech models — MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash — now available in public preview on Azure. The release pushes Microsoft into the top tier of enterprise speech infrastructure at a moment when real-time transcription and voice synthesis are becoming core plumbing for AI assistants, call centers, and productivity tools.
Background
Speech AI has been a fiercely competitive space for the past two years. OpenAI's Whisper established a strong accuracy baseline for transcription; Google and AWS have offered competing cloud speech APIs for years. But enterprise customers have increasingly demanded models that work accurately across many languages simultaneously — not just English — with low enough latency to run in real-time conversations. Microsoft's release today is a direct answer to that demand.
What Microsoft Released
MAI-Transcribe-2-Streaming is the flagship. According to Microsoft's announcement, it supports 60 languages with a 2.5% word error rate (WER) — meaning roughly 2.5 out of every 100 spoken words are transcribed incorrectly. That's a strong accuracy figure at scale. The model ranked #1 on Artificial Analysis' AA-WER Streaming benchmark, which independently measures real-time transcription accuracy across languages. The "streaming" designation means the model returns transcribed text as audio arrives, not after the speaker finishes — a requirement for live captioning, live translation, and real-time AI agents.
MAI-Voice-2.1 and MAI-Voice-2.1-Flash are text-to-speech (TTS) models. The Flash variant is the faster, lower-latency version — suited for conversational AI where response speed matters more than audio fidelity. MAI-Voice-2.1 targets higher-quality voice synthesis for applications where naturalness is the priority.