Alibaba Tongyi Lab · 2026-07-20 · major
Qwen-Audio-3.0-TTS — Alibaba's TTS hits #1 on Artificial Analysis
Qwen-Audio-3.0-TTS is a hosted text-to-speech model from Alibaba's Tongyi Lab, priced at $27.59 per 1M characters and covering 16 languages plus 20 Chinese dialect regions. The Plus tier ranks #1 on the Artificial Analysis speech leaderboard.

Alibaba's new hosted TTS ships in Flash and Plus tiers, spans 16 languages, and takes #1 on the Artificial Analysis speech leaderboard.
Quick facts
| Maker | Alibaba Tongyi Lab (FunAudioLLM) |
|---|---|
| Tiers | Flash (real-time) and Plus (quality) |
| Latency (Flash) | ~300 ms first-packet |
| Price | $27.59 / 1M characters |
| Languages | 16 + 20 Chinese dialect regions |
| Ranking | #1 Artificial Analysis (~1,236 Elo) |
| Access | Hosted via Alibaba Cloud Model Studio |
What is it?
Qwen-Audio-3.0-TTS is a hosted text-to-speech model launched by Alibaba's Tongyi Lab on July 20, 2026, split into Flash (real-time, ~300 ms first-packet latency) and Plus (quality-first). Coverage spans 16 languages plus 20 Chinese dialect regions, all through Alibaba Cloud Model Studio.
How does it work?
The model runs on a 12.5 Hz low-frame-rate speech tokenizer that keeps output legible while cutting per-second compute, and drives synthesis with 86 fine-grained inline tags that let a caller nudge emotion, prosody, and non-verbal cues at the phrase level. Zero-shot voice cloning works from a single reference clip, staying stable through noisy or reverberant reference audio.
Why does it matter?
The Plus tier takes #1 on the Artificial Analysis speech leaderboard at roughly 1,236 Elo, with best-in-class WER or CER in 10 of the 16 supported languages. Combined with $27.59 per 1M characters and a WebSocket streaming API that outputs 48 kHz PCM, WAV, MP3, or Opus, that puts Qwen-Audio-3.0-TTS in direct competition with ElevenLabs and OpenAI Voice on both quality and price.
Who is it for?
Voice agent builders, dubbing studios, and app teams needing multilingual real-time speech
Frequently asked questions
- How much does Qwen-Audio-3.0-TTS cost?
- Qwen-Audio-3.0-TTS is priced at $27.59 per 1 million characters across both tiers on Alibaba Cloud Model Studio. Flash targets real-time interaction with about 300 ms first-packet latency; Plus targets quality-first synthesis and takes the top spot on the Artificial Analysis TTS leaderboard at roughly 1,236 Elo.
- Which languages does Qwen-Audio-3.0-TTS support?
- Qwen-Audio-3.0-TTS covers 16 languages — Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese — plus 20 Chinese dialect regions. Alibaba reports best-in-class WER or CER in 10 of those 16 languages.
- Can Qwen-Audio-3.0-TTS clone a voice?
- Yes. Qwen-Audio-3.0-TTS supports zero-shot voice cloning across all supported languages from a single reference clip, and stays stable when that clip is noisy, reverberant, or unclear. Speaker similarity averages 82.75 for the Plus tier and 80.44 for Flash on Alibaba's internal evaluations.
- Are Qwen-Audio-3.0-TTS weights available?
- No. Qwen-Audio-3.0-TTS is hosted-only through Alibaba Cloud Model Studio, with SDKs and a bidirectional WebSocket streaming API in Singapore and Beijing regions. Weights are not published; buyers get API access at 48 kHz output in PCM, WAV, MP3, or Opus rather than a downloadable checkpoint.
Try it
https://funaudiollm.github.io/qwen-audio-3.0-tts/