October 1, 2026
More accurate streaming speech recognition: MAI-Transcribe-2-Streaming tops the AA-WER ranking
Microsoft ranked first for streaming speech recognition accuracy, but set the price at $0.54 per hour of audio.

Artificial Analysis
@artificialanlys
Microsoft AI released MAI-Transcribe-2-Streaming: the model ranked first for final and first interim transcript accuracy in AA-WER Streaming, with a WER of 2.5% and a latency of 0.13 s after speech ends. MAI-Transcribe-2-Streaming is a new streaming speech recognition model from @MicrosoftAI. It complements MAI-Transcribe-2, which operates without streaming. The model leads in final transcript WER for streaming recognition at 2.5%, ahead of the previous leader, SpaceXAI's Grok Voice Transcribe 2.0 at 2.7%. It returns the final text in 0.13 s instead of 0.49 s. Streaming recognition is available at $0.54 per hour of audio, one of the higher prices among leading streaming models. Highlights ➤ Final transcript: MAI-Transcribe-2-Streaming achieved a WER of 2.5% with a latency of 0.13 s after speech ends, ranking first among 38 models. It is more accurate and faster than Grok Voice Transcribe 2.0 at 2.7% and 0.49 s, Muse Voice Transcribe at 3.1% and 0.16 s, and ElevenLabs Scribe v2 Realtime at 3.6% and 0.14 s. It is also more accurate than Cartesia Ink Preview via external endpoints at 3.1% and 0.11 s, though slightly slower. ➤ First interim transcript: the model achieved a WER of 2.5% with a latency of 0.12 s, ahead of Grok Voice Transcribe 2.0 at 3.4% and 0.49 s. It also beats Muse Voice Transcribe and ElevenLabs Scribe v2 Realtime on accuracy: both have a WER of 3.6%. Microsoft is slightly faster than both: 0.12 s versus 0.13 s. It is more accurate but slower than Cartesia Ink-2 via external endpoints at 4.0% and 0.07 s. ➤ Price: MAI-Transcribe-2-Streaming costs $0.54 per hour of streaming recognition, or $9.00 per 1,000 minutes. That matches Gemini 3.5 Transcribe Live at $9 and is higher than ElevenLabs Scribe v2 Realtime and Deepgram Flux, which charge $6.50. The price is more than twice that of Cartesia Ink-2 at $4 and three times that of Muse Voice Transcribe at $3. Non-streaming recognition costs $0.10 per hour, or $1.67 per 1,000 minutes. Congratulations to the @MicrosoftAI team on the launch! More details below
· 8.1K views
MAI-Transcribe-2-Streaming ranked first among 38 models for final transcript accuracy in AA-WER Streaming as of Oct 1, 2026.
The model converts speech to text as audio arrives. In the Artificial Analysis test, it leads in accuracy for both the first interim transcript and the final transcript. WER measures the word error rate: the lower the score, the more accurate the text.
The previous leader falls behind. Grok Voice Transcribe 2.0 previously topped the final transcript ranking. Microsoft reduced the error rate and shortened the wait for text after speech ends.
| Model | Final transcript error rate | Wait after speech ends | | --- | --- | --- | | MAI-Transcribe-2-Streaming | 2.5% | 0.13 s | | Grok Voice Transcribe 2.0 | 2.7% | 0.49 s | | Muse Voice Transcribe | 3.1% | 0.16 s | | ElevenLabs Scribe v2 Realtime | 3.6% | 0.14 s | | Cartesia Ink Preview, external endpoints | 3.1% | 0.11 s |
Interim text also matters for a voice interface: the application receives it before the final transcript. Microsoft's interim error rate is 2.5%, with a latency of 0.12 s.
The streaming version costs the same as Gemini 3.5 Transcribe Live. The prices below cover the same amount of audio, based on Artificial Analysis data as of Oct 1, 2026.
| Model | Price per 1,000 minutes | | --- | --- | | MAI-Transcribe-2-Streaming | $9 | | Gemini 3.5 Transcribe Live | $9 | | ElevenLabs Scribe v2 Realtime | $6.50 | | Deepgram Flux | $6.50 | | Cartesia Ink-2 | $4 | | Muse Voice Transcribe | $3 |
Connecting in Voice Live. MAI-Transcribe-2 is already available in preview in Azure Voice Live. According to Microsoft Learn documentation dated Sep 24, 2026, you enable it with a `session.update` message by setting `session.input_audio_transcription.model` to `"mai-transcribe-2"`.
Voice Live accepts WebSocket connections at `wss://<your-ai-foundry-resource-name>.services.ai.azure.com/voice-live/realtime?api-version=2026-04-10`. Authentication supports a Microsoft Entra Bearer token or an `api-key`.
In Voice Live, the model is compatible with `gpt-realtime`, `gpt-realtime-mini`, all non-multimodal models and agents. The `input_audio_transcription.phrase_list` field provides hints for recognizing terms, such as your product names.
Primary sources: Artificial Analysis, Oct 1, 2026. Microsoft Learn, Sep 24, 2026.
MAI-Transcribe-2 in Voice Live supports Kazakh, while MAI-Transcribe-1.5 does not.

