AI/TLDR

Alibaba Tongyi Lab · 2026-07-20 · major

Qwen-Audio-3.0-TTS — Alibaba's TTS hits #1 on Artificial Analysis

Qwen-Audio-3.0-TTS is a hosted text-to-speech model from Alibaba's Tongyi Lab, priced at $27.59 per 1M characters and covering 16 languages plus 20 Chinese dialect regions. The Plus tier ranks #1 on the Artificial Analysis speech leaderboard.

Qwen-Audio-3.0-TTS announcement image showing Alibaba Tongyi Lab branding
MarkTechPost

Alibaba's new hosted TTS ships in Flash and Plus tiers, spans 16 languages, and takes #1 on the Artificial Analysis speech leaderboard.

Quick facts

MakerAlibaba Tongyi Lab (FunAudioLLM)
TiersFlash (real-time) and Plus (quality)
Latency (Flash)~300 ms first-packet
Price$27.59 / 1M characters
Languages16 + 20 Chinese dialect regions
Ranking#1 Artificial Analysis (~1,236 Elo)
AccessHosted via Alibaba Cloud Model Studio

What is it?

Qwen-Audio-3.0-TTS is a hosted text-to-speech model launched by Alibaba's Tongyi Lab on July 20, 2026, split into Flash (real-time, ~300 ms first-packet latency) and Plus (quality-first). Coverage spans 16 languages plus 20 Chinese dialect regions, all through Alibaba Cloud Model Studio.

How does it work?

The model runs on a 12.5 Hz low-frame-rate speech tokenizer that keeps output legible while cutting per-second compute, and drives synthesis with 86 fine-grained inline tags that let a caller nudge emotion, prosody, and non-verbal cues at the phrase level. Zero-shot voice cloning works from a single reference clip, staying stable through noisy or reverberant reference audio.

Why does it matter?

The Plus tier takes #1 on the Artificial Analysis speech leaderboard at roughly 1,236 Elo, with best-in-class WER or CER in 10 of the 16 supported languages. Combined with $27.59 per 1M characters and a WebSocket streaming API that outputs 48 kHz PCM, WAV, MP3, or Opus, that puts Qwen-Audio-3.0-TTS in direct competition with ElevenLabs and OpenAI Voice on both quality and price.

Who is it for?

Voice agent builders, dubbing studios, and app teams needing multilingual real-time speech

Frequently asked questions

How much does Qwen-Audio-3.0-TTS cost?
Qwen-Audio-3.0-TTS is priced at $27.59 per 1 million characters across both tiers on Alibaba Cloud Model Studio. Flash targets real-time interaction with about 300 ms first-packet latency; Plus targets quality-first synthesis and takes the top spot on the Artificial Analysis TTS leaderboard at roughly 1,236 Elo.
Which languages does Qwen-Audio-3.0-TTS support?
Qwen-Audio-3.0-TTS covers 16 languages — Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese — plus 20 Chinese dialect regions. Alibaba reports best-in-class WER or CER in 10 of those 16 languages.
Can Qwen-Audio-3.0-TTS clone a voice?
Yes. Qwen-Audio-3.0-TTS supports zero-shot voice cloning across all supported languages from a single reference clip, and stays stable when that clip is noisy, reverberant, or unclear. Speaker similarity averages 82.75 for the Plus tier and 80.44 for Flash on Alibaba's internal evaluations.
Are Qwen-Audio-3.0-TTS weights available?
No. Qwen-Audio-3.0-TTS is hosted-only through Alibaba Cloud Model Studio, with SDKs and a bidirectional WebSocket streaming API in Singapore and Beijing regions. Weights are not published; buyers get API access at 48 kHz output in PCM, WAV, MP3, or Opus rather than a downloadable checkpoint.

Try it

https://funaudiollm.github.io/qwen-audio-3.0-tts/

Sources · 2 outlets

Tags

  • qwen
  • qwen-audio
  • alibaba
  • tongyi-lab
  • tts
  • text-to-speech
  • voice-cloning
  • multilingual
  • streaming

← All releases · Learn AI