Microsoft launches MAI streaming transcription and two Voice 2.1 models
Microsoft launched MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash for developers building real-time multilingual voice applications.
Microsoft launched MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash on October 1, bringing real-time speech recognition and two speech-generation options to developers building conversational voice applications.
Microsoft says developers can access all three models through Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live. The voice models are also available through OpenRouter, while LiveKit support is listed as coming soon. Azure documentation marks the streaming model and both Voice 2.1 variants as public previews. The releases extend Microsoft’s Foundry voice-agent toolset.
MAI-Transcribe-2-Streaming accepts continuous audio, returning partial transcripts that update before each segment is finalized. It supports an OpenAI Realtime-compatible WebSocket interface and the Azure Speech SDK. Microsoft says the model covers 60 languages, detects language automatically and continuously, and starts returning partial hypotheses just over 100 milliseconds after receiving audio. The latency figure has not been independently reproduced.
Microsoft lists an introductory price of $0.54 per hour of audio through the end of 2026. It also says the model ranks first for both partial- and final-transcript accuracy on Artificial Analysis. That ranking could not be independently confirmed because the accessible benchmark page did not show model-level results.
For speech generation, Microsoft positions MAI-Voice-2.1 for high-fidelity, longer-form output, while MAI-Voice-2.1-Flash targets latency-sensitive, high-volume agents and assistants. Both support 23 languages. Microsoft says Voice 2.1 covers 26 locales and can maintain one speaker identity across languages. The listed prices are $22 per million characters for Voice 2.1 and $15 per million characters for Voice 2.1-Flash.
Microsoft claims the Flash model can generate 45 seconds of audio with 150 milliseconds of end-to-end latency. The company also says it cuts model inference time by 55% and costs about 60% less than comparable models. Those comparative claims have not been independently reproduced. Both voice models offer gated instant voice cloning from a consented reference clip. Microsoft documentation specifies a clip lasting five to 60 seconds and says consent guardrails are built in.
The October streaming release is distinct from MAI-Transcribe-2, which Microsoft announced on September 3. Microsoft presented that earlier model for batch transcription, with features including diarization, word-level timestamps, keyword biasing and configurable transcription styles. The new model instead processes a continuous audio stream, updating its transcript while a speaker is talking.
More news

Databricks adds branch-based restores to Lakebase Postgres

AWS adds hierarchy filtering to Amazon Quick Sight dashboards

ServiceNow CoreAI introduces AutoSynthData for enterprise-agent training
