█

AI/TLDR

MAI-Transcribe-2-Streaming

Microsoft AI's first streaming speech-to-text model, announced 1 October 2026 — real-time transcripts in 60 languages with first partials in just over 100ms.

MAI-Transcribe (speech-to-text)API onlyPreview
Released
1 Oct 2026
Input
$0.54 / hour of audio
License
Proprietary

Overview

MAI-Transcribe-2-Streaming is the real-time speech-to-text model Microsoft AI announced on 1 October 2026, described in the launch post as its first streaming transcription model. It was released alongside two voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft says it "delivers low-latency, real-time transcripts in 60 languages, all while supporting automatic, continuous language detection."

Instead of waiting for a speaker to finish, the model returns provisional hypotheses ("partials") as audio arrives — Microsoft says the first one lands "in just over 100ms of receiving audio" — and then confirms each segment with a final transcript. For real-time dictation and subtitling, Microsoft says its internal evaluations show words appear in the transcript 2x faster than with its closest competitor.

Microsoft says the model ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis. On the Artificial Analysis AA-WER Streaming leaderboard (read 4 October 2026) it scores 2.51% word error rate, ahead of Grok Voice Transcribe 2.0 (Streaming) at 2.73%, and delivers its final transcript 0.13 seconds after end of speech.

On Microsoft Foundry the model is in public preview: Microsoft Learn documents two integrations — an OpenAI Realtime-compatible WebSocket API and the Azure Speech SDK (v1.52.0) — with each Realtime session lasting up to one hour and raw 16-bit PCM mono audio at 16 kHz or 24 kHz. The launch post lists an introductory price of $0.54 per hour of audio through the end of 2026, and access through Microsoft Foundry, MAI Playground, Vercel and Azure Voice Live, with LiveKit support marked as coming soon.

Released2026-10-01
LicenseProprietary
WeightsAPI only
ModalitiesAudio, Text
StatusPreview

Benchmarks

Artificial Analysis AA-WER Streaming leaderboard, read 4 October 2026 (≈8 hours of audio across AA-AgentTalk, VoxPopuli and Earnings22). Lower is better on every row.

BenchmarkMAI-Transcribe-2-StreamingGrok Voice Transcribe 2.0 (Streaming)ElevenLabs Scribe v2 RealtimeGPT Live TranscribeGemini 3.5 Transcribe Live
AA-WER Streaming index (final transcript)2.51%2.73%3.59%3.92%4%
WER of first partial after end of speech2.52%3.36%3.59%6.35%5.77%
Time to final transcript after end of speech0.13 s0.49 s0.14 s0.81 s0.4 s
Time to first partial transcript0.12 s0.49 s0.13 s0.26 s0.25 s
Price per 1,000 minutes of audio9 $3.33 $6.5 $17 $9 $

Comparison source ↗

Pricing

Input$0.54 / hour of audio

Introductory price through the end of 2026, per Microsoft's launch post.

Pricing source ↗

Strengths

  • Ranks first for final-transcript accuracy on the Artificial Analysis AA-WER Streaming leaderboard, at 2.51% word error rate (read 4 October 2026)
  • First partial transcripts in just over 100ms of receiving audio, per Microsoft
  • Words appear in the transcript 2x faster than with Microsoft's closest competitor for dictation and subtitling, by Microsoft's internal evaluations
  • 60 languages with automatic, continuous language detection — no language needs to be declared up front
  • OpenAI Realtime-compatible WebSocket protocol, plus a managed Azure Speech SDK integration
  • Introductory price of $0.54 per hour of audio through the end of 2026

Best for

  • Reach for it for voice agents that need a transcript while the user is still speaking, the case Microsoft built it for
  • Reach for it for call-centre transcription and voice assistants, workloads Microsoft Learn lists for the model
  • Reach for it for live meeting and lecture captioning or subtitling, where partial text has to appear as people talk
  • Reach for it for real-time dictation and note taking across multilingual speakers, using automatic language detection

How to access

ProviderModel ID
Microsoft Foundry (Realtime API) ↗MAI-Transcribe-2-Streaming
Azure Speech SDK ↗MAI-Transcribe-2-Streaming
Vercel AI Gateway ↗microsoft/mai-transcribe-2-streaming

FAQ

What is MAI-Transcribe-2-Streaming?

It is Microsoft AI's real-time speech-to-text model, announced on 1 October 2026 as its first streaming transcription model. Audio is sent as a continuous stream and the model returns partial transcripts while the speaker talks, then a final transcript for each segment.

How accurate is MAI-Transcribe-2-Streaming?

Microsoft says it ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis. On the AA-WER Streaming leaderboard, read 4 October 2026, it scores 2.51% word error rate, ahead of Grok Voice Transcribe 2.0 (Streaming) at 2.73%.

How fast is it?

Microsoft says the first partial transcript arrives in just over 100ms of receiving audio, and that its internal evaluations show words appearing 2x faster than with its closest competitor for dictation and subtitling. Artificial Analysis measures 0.13 seconds from end of speech to the final transcript.

What does MAI-Transcribe-2-Streaming cost?

Microsoft's launch post lists an introductory price of $0.54 per hour of audio through the end of 2026.

How many languages does it support?

60 languages, with automatic and continuous language detection, so a session does not have to declare its language up front. A language code can optionally be passed as a hint.

How do I access it?

Through Microsoft Foundry — via an OpenAI Realtime-compatible WebSocket API or the Azure Speech SDK (v1.52.0) — as well as MAI Playground, Vercel AI Gateway (`microsoft/mai-transcribe-2-streaming`) and Azure Voice Live. Microsoft lists LiveKit support as coming soon. On Foundry the model is documented as a public preview without a service-level agreement.

Does it return word timestamps or speaker labels?

Not through the Speech SDK: Microsoft Learn states results don't include detected-language information, confidence scores or word-level timestamps, and only final results carry segment-level offset and duration. The documentation does not list speaker diarization for the model.