Alibaba Qwen · 2026-09-23 · major
Qwen-Audio-3.1 — five speech models and API price cuts of up to 95%
Alibaba's Qwen team released Qwen-Audio-3.1 on 23 September 2026: upgraded ASR, TTS and Realtime models plus new ASR-Next and TTS-Next. API prices drop about 70% for TTS, 85% for Realtime and up to 95% for ASR.

Alibaba's Qwen team ships a full speech stack — hear, speak, talk live, and make sound — and cuts its voice API prices sharply.
Quick facts
| Maker | Alibaba Qwen team |
|---|---|
| Models | ASR, ASR-Next, TTS, TTS-Next, Realtime |
| ASR languages | 30 languages + 16 Chinese dialects |
| ASR latency | ~160 ms to first character |
| Price cuts | TTS ~70%, Realtime ~85%, ASR up to 95% |
| Availability | Hosted API (Alibaba Cloud Model Studio) |
What is it?
Qwen-Audio-3.1 is a family of five hosted speech models from Alibaba's Qwen team. Three are upgrades of existing models: ASR for speech recognition, TTS for speech synthesis, and Realtime for live voice conversation. Two are new: ASR-Next, which understands audio beyond words, and TTS-Next, which creates whole audio scenes.
How does it work?
ASR-Next labels different speakers with timestamps and also detects emotions, background sounds and machine noise. TTS-Next pairs a language model with a diffusion approach, so voice, sound effects and ambient audio come out in a single pass from text, timestamps and up to three reference clips. The Realtime model listens and speaks at the same time and can be interrupted mid-sentence.
Why does it matter?
A 70–95% cut in API cost lowers the bill for any voice agent or transcription app built on Qwen, and one vendor now covers hearing, speaking, live talk and sound design. Alibaba reports an average 4.55% character error rate for the ASR model on public dialect tests, ahead of rivals on six of eleven subsets.
Who is it for?
developers building voice agents, transcription and audio content
Frequently asked questions
- Are the Qwen-Audio-3.1 models open weights?
- Qwen-Audio-3.1 is offered as hosted APIs through Alibaba Cloud Model Studio, where each model has its own documentation and price page. The coverage of the 23 September 2026 launch describes API access and price cuts, and does not mention downloadable weights, so plan on calling it as a cloud service rather than running it locally.
- What can Qwen-Audio-3.1-TTS-Next generate?
- Qwen-Audio-3.1-TTS-Next generates complete audio in one pass, mixing speech, sound effects and ambient sound. It accepts up to 3,000 characters of text plus up to three 30-second reference clips, and outputs WAV, MP3 or PCM of up to 120 seconds, or 240 seconds for podcasts. The model documentation lists Chinese and English support.
- How much does Qwen-Audio-3.1-TTS-Next cost?
- Alibaba Cloud Model Studio lists Qwen-Audio-3.1-TTS-Next in the China (Beijing) region at $0.848 per million input tokens and $1.696 per million output tokens, with a limit of 3 requests per second. Prices in other regions and for the other four Qwen-Audio-3.1 models are listed on their own Model Studio pages.
- What is new in Qwen-Audio-3.1-ASR-Next compared with plain ASR?
- Qwen-Audio-3.1-ASR-Next goes beyond a transcript. It separates speakers with timestamps and aligned text, and it recognises emotions, ambient sounds and machine noise, which supports tasks such as sound captioning, event localisation and questions about an audio clip. The standard Qwen-Audio-3.1-ASR focuses on accurate multilingual and dialect transcription with filler words removed.
Try it
Model id qwen-audio-3.1-tts-next in Alibaba Cloud Model Studio