█

AI/TLDR

MAI-Voice-2.1-Flash

Microsoft AI's low-latency text-to-speech model, announced 1 October 2026 — 45 seconds of audio with 150ms end-to-end latency, in 23 languages, for voice agents.

MAI-Voice (text-to-speech)API onlyPreview
Released
1 Oct 2026
Input
$15 / 1M characters
License
Proprietary

Overview

MAI-Voice-2.1-Flash is the fast variant of MAI-Voice-2.1, announced by Microsoft AI on 1 October 2026 together with MAI-Voice-2.1 and the streaming speech-to-text model MAI-Transcribe-2-Streaming. Microsoft says it "supports the same languages, and cross-language speakers, as MAI-Voice-2.1" — 23 languages and 26 locales — "but it's been leveled up for high-volume, latency-sensitive workloads."

The launch post puts numbers on that: the model "can generate 45s of audio, with an end-to-end latency, of a mere 150ms", and Microsoft says it delivers 55% faster model inference and is about 60% cheaper than comparable models. Microsoft Learn positions it for real-time voice agents and assistants, low-latency call-centre and IVR flows, and multilingual interactive experiences.

Microsoft pitches it as the speaking half of a voice agent: "Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends" — the streaming model transcribes the caller quickly and Flash answers at conversational speed. Like MAI-Voice-2.1, it supports fine-grained emotion and style control through SSML `mstts:express-as`, a library of licensed prebuilt voices, and gated instant voice cloning from a 5–60 second consented reference clip.

On Microsoft Foundry the model is in public preview, served globally from 14 Azure regions. It is called through the Azure Speech SDK with SSML — a prebuilt voice takes the model as a suffix, such as `en-US-Harper:MAI-Voice-2.1-Flash` — and through Azure Voice Live. Microsoft's launch post lists a price of $15 per 1M characters and access through Microsoft Foundry, MAI Playground, OpenRouter, Vercel and Azure Voice Live, with LiveKit support marked as coming soon.

Released2026-10-01
LicenseProprietary
WeightsAPI only
ModalitiesText, Audio
StatusPreview

Pricing

Input$15 / 1M characters

Price listed in Microsoft AI's launch post of 1 October 2026.

Pricing source ↗

Strengths

  • Generates 45 seconds of audio with 150ms end-to-end latency, per Microsoft's launch post
  • 55% faster model inference, by Microsoft's account
  • $15 per 1M characters, which Microsoft says is about 60% cheaper than comparable models
  • Same 23 languages, 26 locales and cross-language voices as MAI-Voice-2.1
  • Emotion and style control through SSML `mstts:express-as`, plus gated instant voice cloning from a 5–60 second reference clip
  • Available in Azure Voice Live as well as the Azure Speech SDK, served from 14 Azure regions

Best for

  • Reach for it for real-time voice agents and assistants, the workload Microsoft built it for
  • Reach for it for low-latency call-centre and IVR flows, a use Microsoft Learn lists for the model
  • Reach for it as the speech output of a voice agent paired with MAI-Transcribe-2-Streaming, as Microsoft recommends
  • Reach for MAI-Voice-2.1 instead for audiobooks and long-form narration, where fidelity matters more than speed

How to access

ProviderModel ID
Microsoft Foundry (Azure Speech) ↗MAI-Voice-2.1-Flash
OpenRouter ↗microsoft/mai-voice-2.1-flash
Vercel AI Gateway ↗microsoft/mai-voice-2.1-flash

MAI-Voice (text-to-speech) — every version

The full lineage of the MAI-Voice (text-to-speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
MAI-Voice-2.1current2026-10-01—Proprietary
MAI-Voice-2.1-Flash2026-10-01—Proprietary

FAQ

What is MAI-Voice-2.1-Flash?

It is the fast, low-latency variant of Microsoft AI's MAI-Voice-2.1 text-to-speech model, announced on 1 October 2026 for high-volume, latency-sensitive workloads such as voice agents and call-centre flows.

How fast is MAI-Voice-2.1-Flash?

Microsoft says it can generate 45 seconds of audio with an end-to-end latency of 150ms, and that it delivers 55% faster model inference.

What does MAI-Voice-2.1-Flash cost?

Microsoft's launch post lists $15 per 1M characters, which it describes as about 60% cheaper than comparable models. MAI-Voice-2.1 is listed at $22 per 1M characters.

Does the Flash model support fewer languages than MAI-Voice-2.1?

No. Microsoft says it supports the same languages and cross-language speakers as MAI-Voice-2.1 — 23 languages and 26 locales — and every managed voice in the Microsoft Learn voice table supports both models.

Can it clone a voice?

Yes, with gated access. Instant voice cloning uses a short consented reference clip (Microsoft Learn recommends 5–60 seconds); you apply through Microsoft's Limited Access Review, and only authorised, licensed voices can be synthesised.

How do I use MAI-Voice-2.1-Flash?

On Microsoft Foundry it is in public preview: call it through the Azure Speech SDK or REST API with SSML and a voice name such as `en-US-Harper:MAI-Voice-2.1-Flash`, or through Azure Voice Live. Microsoft also lists MAI Playground, OpenRouter and Vercel AI Gateway (`microsoft/mai-voice-2.1-flash`), with LiveKit support coming soon.