01 What happened
Microsoft announced three MAI models in Microsoft Foundry: MAI-Transcribe-2-Streaming for streaming speech transcription, MAI-Voice-2.1 for multilingual text-to-speech, and MAI-Voice-2.1-Flash for responsive, high-volume voice applications. All three are available through Azure Speech and direct API access through Microsoft Foundry.
02 Key details
- MAI-Transcribe-2-Streaming transcribes continuously across 60 languages with automatic language detection. It emits first partial hypotheses within the low hundreds of milliseconds after receiving audio and commits a stable transcript once the utterance ends.
- Microsoft says the model debuted at #1 for partial- and final-transcript accuracy on the Artificial Analysis leaderboard. In most cases, words appear in the transcript as early as 320 ms after being spoken, compared with more than 500 ms for the closest competition.
- MAI-Voice-2.1 generates speech in 23 languages with a consistent voice identity across supported languages. In a supplied comparison Microsoft attributes to ElevenLabs documentation, MAI-Voice-2.1-Flash is 55% faster and 60% less expensive than competing models in its class.
- MAI-Transcribe-2-Streaming has an introductory price of $0.54 per hour of audio through the end of the year. MAI-Voice-2.1 costs $22 per 1M characters and MAI-Voice-2.1-Flash costs $15 per 1M characters. OpenRouter, Vercel, and LiveKit are named only as planned platforms.
03 Why it matters
Developers can pair streaming transcription, with first partials arriving within the low hundreds of milliseconds, with either of two text-to-speech models. Published per-hour and per-1M-character prices support cost planning. Microsoft’s performance claims have not been independently verified here.
04 Who it matters to
Voice interface developers, service application developers, AI agent developers and teams selecting TTS models.