█

AI/TLDR

Alibaba Qwen · 2026-09-23 · major

Qwen-Audio-3.1 — five speech models and API price cuts of up to 95%

Alibaba's Qwen team released Qwen-Audio-3.1 on 23 September 2026: upgraded ASR, TTS and Realtime models plus new ASR-Next and TTS-Next. API prices drop about 70% for TTS, 85% for Realtime and up to 95% for ASR.

Qwen Audio 3.1 stack graphic covering speech recognition, text-to-speech and realtime voice
tbreak

Alibaba's Qwen team ships a full speech stack — hear, speak, talk live, and make sound — and cuts its voice API prices sharply.

Quick facts

MakerAlibaba Qwen team
ModelsASR, ASR-Next, TTS, TTS-Next, Realtime
ASR languages30 languages + 16 Chinese dialects
ASR latency~160 ms to first character
Price cutsTTS ~70%, Realtime ~85%, ASR up to 95%
AvailabilityHosted API (Alibaba Cloud Model Studio)

What is it?

Qwen-Audio-3.1 is a family of five hosted speech models from Alibaba's Qwen team. Three are upgrades of existing models: ASR for speech recognition, TTS for speech synthesis, and Realtime for live voice conversation. Two are new: ASR-Next, which understands audio beyond words, and TTS-Next, which creates whole audio scenes.

How does it work?

ASR-Next labels different speakers with timestamps and also detects emotions, background sounds and machine noise. TTS-Next pairs a language model with a diffusion approach, so voice, sound effects and ambient audio come out in a single pass from text, timestamps and up to three reference clips. The Realtime model listens and speaks at the same time and can be interrupted mid-sentence.

Why does it matter?

A 70–95% cut in API cost lowers the bill for any voice agent or transcription app built on Qwen, and one vendor now covers hearing, speaking, live talk and sound design. Alibaba reports an average 4.55% character error rate for the ASR model on public dialect tests, ahead of rivals on six of eleven subsets.

Who is it for?

developers building voice agents, transcription and audio content

Frequently asked questions

Are the Qwen-Audio-3.1 models open weights?
Qwen-Audio-3.1 is offered as hosted APIs through Alibaba Cloud Model Studio, where each model has its own documentation and price page. The coverage of the 23 September 2026 launch describes API access and price cuts, and does not mention downloadable weights, so plan on calling it as a cloud service rather than running it locally.
What can Qwen-Audio-3.1-TTS-Next generate?
Qwen-Audio-3.1-TTS-Next generates complete audio in one pass, mixing speech, sound effects and ambient sound. It accepts up to 3,000 characters of text plus up to three 30-second reference clips, and outputs WAV, MP3 or PCM of up to 120 seconds, or 240 seconds for podcasts. The model documentation lists Chinese and English support.
How much does Qwen-Audio-3.1-TTS-Next cost?
Alibaba Cloud Model Studio lists Qwen-Audio-3.1-TTS-Next in the China (Beijing) region at $0.848 per million input tokens and $1.696 per million output tokens, with a limit of 3 requests per second. Prices in other regions and for the other four Qwen-Audio-3.1 models are listed on their own Model Studio pages.
What is new in Qwen-Audio-3.1-ASR-Next compared with plain ASR?
Qwen-Audio-3.1-ASR-Next goes beyond a transcript. It separates speakers with timestamps and aligned text, and it recognises emotions, ambient sounds and machine noise, which supports tasks such as sound captioning, event localisation and questions about an audio clip. The standard Qwen-Audio-3.1-ASR focuses on accurate multilingual and dialect transcription with filler words removed.

Try it

Model id qwen-audio-3.1-tts-next in Alibaba Cloud Model Studio

Sources · 4 outlets

Tags

  • qwen
  • qwen-audio
  • alibaba
  • speech
  • asr
  • tts
  • text-to-speech
  • speech-recognition
  • realtime-voice
  • voice-agents
  • pricing
  • model-studio

← All releases · Learn AI