█

AI/TLDR

MAI-Voice-2.1

Microsoft AI's high-fidelity text-to-speech model, announced 1 October 2026 — expressive speech in 23 languages and 26 locales, with one voice keeping its identity across every language.

MAI-Voice (text-to-speech)API onlyPreview
Released
1 Oct 2026
Input
$22 / 1M characters
License
Proprietary

Overview

MAI-Voice-2.1 is the text-to-speech model Microsoft AI announced on 1 October 2026, alongside a faster sibling, MAI-Voice-2.1-Flash, and the streaming speech-to-text model MAI-Transcribe-2-Streaming. Microsoft calls it "our strongest multilingual text-to-speech model yet", and says it has been expanded to support 23 languages and 26 locales "while enabling one single 'voice' to use all languages with a truly native accent."

Microsoft Learn describes it as the highest-fidelity, most expressive model in the MAI-Voice family, built for experiences where maximum voice quality matters, such as long-form narration and brand-defining audio. It is optimised for long-form generation with stable persona quality and speaker consistency across audiobooks, podcasts and lectures. The documentation is explicit about the trade-off: the model prioritises naturalness and expressivity over latency-critical scenarios, which is the job of MAI-Voice-2.1-Flash.

Speaking style is controlled through SSML: the `mstts:express-as` element with a `style` attribute steers emotions such as joy, excitement and empathy, and the prebuilt voices in the managed voice library each list their supported styles (for example `audiobook`, `narrator`, `whispering` or `sad`). Both MAI-Voice-2.1 models also support instant voice cloning from a short reference clip — Microsoft Learn recommends 5–60 seconds of audio — across all supported languages, but the feature is gated: only authorised, consented voices can be synthesised, and access requires approval through Microsoft's Limited Access Review.

On Microsoft Foundry the model is in public preview, served globally from 14 Azure regions through the same Azure Speech API and SDK as other Azure neural voices: a prebuilt voice name takes the model as a suffix, such as `en-US-Harper:MAI-Voice-2.1`. Microsoft's launch post lists a price of $22 per 1M characters and access through Microsoft Foundry, MAI Playground, OpenRouter, Vercel and Azure Voice Live, with LiveKit support marked as coming soon.

Released2026-10-01
LicenseProprietary
WeightsAPI only
ModalitiesText, Audio
StatusPreview

Pricing

Input$22 / 1M characters

Price listed in Microsoft AI's launch post of 1 October 2026.

Pricing source ↗

Strengths

  • 23 languages and 26 locales, with one voice able to speak every supported language with a native accent, per Microsoft
  • Highest-fidelity, most expressive model in the MAI-Voice family, according to Microsoft Learn
  • Optimised for long-form generation with stable persona quality and speaker consistency across extended content
  • Fine-grained emotion and style control through SSML `mstts:express-as`
  • Gated instant voice cloning from a 5–60 second reference clip, with built-in consent guardrails
  • Uses the standard Azure Speech REST API and SDKs (Python, C#, JavaScript, Java), served from 14 Azure regions

Best for

  • Reach for it for audiobooks and podcasts, where Microsoft Learn says it keeps a speaker consistent across long-form content
  • Reach for it for educational content and lectures that need expressive narration in several of its 23 languages
  • Reach for it for voice-overs and brand-defining audio, including a consented clone of a brand's own voice
  • Reach for MAI-Voice-2.1-Flash instead when latency matters more than maximum fidelity, such as live voice agents

How to access

ProviderModel ID
Microsoft Foundry (Azure Speech) ↗MAI-Voice-2.1
OpenRouter ↗microsoft/mai-voice-2.1
Vercel AI Gateway ↗microsoft/mai-voice-2.1

MAI-Voice (text-to-speech) — every version

The full lineage of the MAI-Voice (text-to-speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
MAI-Voice-2.1current2026-10-01—Proprietary
MAI-Voice-2.1-Flash2026-10-01—Proprietary

FAQ

What is MAI-Voice-2.1?

It is Microsoft AI's text-to-speech model, announced on 1 October 2026. Microsoft Learn describes it as the highest-fidelity, most expressive model in the MAI-Voice family, aimed at long-form narration, audiobooks, podcasts and voice-overs.

How many languages does MAI-Voice-2.1 support?

Microsoft says 23 languages and 26 locales, and that a single voice can speak all of them with a native accent, so a speaker keeps the same identity when switching language.

What does MAI-Voice-2.1 cost?

Microsoft's launch post lists $22 per 1M characters. The faster MAI-Voice-2.1-Flash is listed at $15 per 1M characters.

What is the difference between MAI-Voice-2.1 and MAI-Voice-2.1-Flash?

Both cover the same languages and cross-language voices. MAI-Voice-2.1 prioritises naturalness and expressivity over latency and is optimised for long-form generation; MAI-Voice-2.1-Flash is built for high-volume, latency-sensitive workloads such as voice agents, and is cheaper.

Can MAI-Voice-2.1 clone a voice?

Yes, but access is gated. Instant voice cloning works from a short reference clip (Microsoft Learn recommends 5–60 seconds) across all supported languages; you must apply through Microsoft's Limited Access Review, upload audio consent, and only authorised, licensed voices can be synthesised.

How do I use MAI-Voice-2.1?

On Microsoft Foundry it is in public preview and works through the standard Azure Speech REST API and SDKs: send SSML with a voice name such as `en-US-Harper:MAI-Voice-2.1`. Microsoft also lists MAI Playground, OpenRouter and Vercel AI Gateway (`microsoft/mai-voice-2.1`) and Azure Voice Live, with LiveKit support coming soon.