Overview
MAI-Voice-2.1-Flash is the fast variant of MAI-Voice-2.1, announced by Microsoft AI on 1 October 2026 together with MAI-Voice-2.1 and the streaming speech-to-text model MAI-Transcribe-2-Streaming. Microsoft says it "supports the same languages, and cross-language speakers, as MAI-Voice-2.1" — 23 languages and 26 locales — "but it's been leveled up for high-volume, latency-sensitive workloads."
The launch post puts numbers on that: the model "can generate 45s of audio, with an end-to-end latency, of a mere 150ms", and Microsoft says it delivers 55% faster model inference and is about 60% cheaper than comparable models. Microsoft Learn positions it for real-time voice agents and assistants, low-latency call-centre and IVR flows, and multilingual interactive experiences.
Microsoft pitches it as the speaking half of a voice agent: "Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends" — the streaming model transcribes the caller quickly and Flash answers at conversational speed. Like MAI-Voice-2.1, it supports fine-grained emotion and style control through SSML `mstts:express-as`, a library of licensed prebuilt voices, and gated instant voice cloning from a 5–60 second consented reference clip.
On Microsoft Foundry the model is in public preview, served globally from 14 Azure regions. It is called through the Azure Speech SDK with SSML — a prebuilt voice takes the model as a suffix, such as `en-US-Harper:MAI-Voice-2.1-Flash` — and through Azure Voice Live. Microsoft's launch post lists a price of $15 per 1M characters and access through Microsoft Foundry, MAI Playground, OpenRouter, Vercel and Azure Voice Live, with LiveKit support marked as coming soon.
| Released | 2026-10-01 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Modalities | Text, Audio |
| Status | Preview |
Pricing
| Input | $15 / 1M characters |
|---|
Price listed in Microsoft AI's launch post of 1 October 2026.
Strengths
- Generates 45 seconds of audio with 150ms end-to-end latency, per Microsoft's launch post
- 55% faster model inference, by Microsoft's account
- $15 per 1M characters, which Microsoft says is about 60% cheaper than comparable models
- Same 23 languages, 26 locales and cross-language voices as MAI-Voice-2.1
- Emotion and style control through SSML `mstts:express-as`, plus gated instant voice cloning from a 5–60 second reference clip
- Available in Azure Voice Live as well as the Azure Speech SDK, served from 14 Azure regions
Best for
- Reach for it for real-time voice agents and assistants, the workload Microsoft built it for
- Reach for it for low-latency call-centre and IVR flows, a use Microsoft Learn lists for the model
- Reach for it as the speech output of a voice agent paired with MAI-Transcribe-2-Streaming, as Microsoft recommends
- Reach for MAI-Voice-2.1 instead for audiobooks and long-form narration, where fidelity matters more than speed
How to access
| Provider | Model ID |
|---|---|
| Microsoft Foundry (Azure Speech) ↗ | MAI-Voice-2.1-Flash |
| OpenRouter ↗ | microsoft/mai-voice-2.1-flash |
| Vercel AI Gateway ↗ | microsoft/mai-voice-2.1-flash |
MAI-Voice (text-to-speech) — every version
The full lineage of the MAI-Voice (text-to-speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| MAI-Voice-2.1current | 2026-10-01 | — | Proprietary |
| MAI-Voice-2.1-Flash | 2026-10-01 | — | Proprietary |
FAQ
What is MAI-Voice-2.1-Flash?
It is the fast, low-latency variant of Microsoft AI's MAI-Voice-2.1 text-to-speech model, announced on 1 October 2026 for high-volume, latency-sensitive workloads such as voice agents and call-centre flows.
How fast is MAI-Voice-2.1-Flash?
Microsoft says it can generate 45 seconds of audio with an end-to-end latency of 150ms, and that it delivers 55% faster model inference.
What does MAI-Voice-2.1-Flash cost?
Microsoft's launch post lists $15 per 1M characters, which it describes as about 60% cheaper than comparable models. MAI-Voice-2.1 is listed at $22 per 1M characters.
Does the Flash model support fewer languages than MAI-Voice-2.1?
No. Microsoft says it supports the same languages and cross-language speakers as MAI-Voice-2.1 — 23 languages and 26 locales — and every managed voice in the Microsoft Learn voice table supports both models.
Can it clone a voice?
Yes, with gated access. Instant voice cloning uses a short consented reference clip (Microsoft Learn recommends 5–60 seconds); you apply through Microsoft's Limited Access Review, and only authorised, licensed voices can be synthesised.
How do I use MAI-Voice-2.1-Flash?
On Microsoft Foundry it is in public preview: call it through the Azure Speech SDK or REST API with SSML and a voice name such as `en-US-Harper:MAI-Voice-2.1-Flash`, or through Azure Voice Live. Microsoft also lists MAI Playground, OpenRouter and Vercel AI Gateway (`microsoft/mai-voice-2.1-flash`), with LiveKit support coming soon.