█

AI/TLDR

Qwen-Audio-3.1-ASR

Alibaba's hosted speech-recognition model from the Qwen-Audio-3.1 release of 23 September 2026 — multilingual and Chinese-dialect transcription with built-in polishing.

Qwen-Audio (speech)API only
Released
23 Sep 2026
Context
8K
Parameters
Not disclosed
Input
$0.15 / 1M tokens
License
Proprietary (hosted API)

Overview

Qwen-Audio-3.1-ASR is the speech-recognition model in Qwen-Audio-3.1, the speech release Alibaba's Qwen team announced on 23 September 2026. The launch post credits it with stronger multilingual and dialect recognition and native polishing that removes filler words and repetitions, so transcripts read as clean text rather than a verbatim record.

Alibaba serves it in three forms. `qwen-audio-3.1-asr-flash` handles short audio over HTTP, with context enhancement, punctuation prediction, text normalization and speaker-role transcription for multi-person dialogue. `qwen-audio-3.1-asr-flash-filetrans` is the offline file version for meeting transcription, content production and call analysis, adding hot words and speaker separation. `qwen-audio-3.1-asr-flash-streaming` is the real-time version for live subtitles and conferencing.

The short-audio model has an 8,192-token context window, with up to 7,168 input and 1,024 output tokens, and a limit of 600 requests per minute. It lists $0.15 per 1M input tokens and $0.47 per 1M output tokens in Singapore, and $0.113 and $0.382 in China (Beijing). The streaming version lists $0.93 and $0.70 in Singapore.

The launch post puts the ASR price cut at up to 95%. A sibling model, Qwen-Audio-3.1-ASR-Next, adds speaker labels with timestamps and recognises emotions, ambient sounds and machine noise for audio captioning and question answering.

Released2026-09-23
LicenseProprietary (hosted API)
WeightsAPI only
ParametersNot disclosed
Context8K
ModalitiesAudio, Text

Pricing

Input$0.15 / 1M tokens
Output$0.47 / 1M tokens

Singapore rates for qwen-audio-3.1-asr-flash; China (Beijing) is $0.113 / $0.382. The streaming model lists $0.93 / $0.70 in Singapore.

Pricing source ↗

Strengths

  • Multilingual recognition with coverage of Chinese dialects from many regions
  • Native polishing that drops filler words and repetitions from the transcript
  • Short-audio, offline-file and real-time streaming versions in one model generation
  • $0.15 / $0.47 per 1M input / output tokens in Singapore for the short-audio model

Best for

  • Reach for it for meeting notes and call analysis, using the file-transcription version with speaker separation
  • Reach for it for live subtitles and voice interfaces, using the streaming version
  • Reach for it for transcripts of Chinese dialect speech that should read cleanly without manual editing
  • Reach for Qwen-Audio-3.1-ASR-Next instead when you also need emotions, sound events or audio question answering

How to access

ProviderModel ID
Alibaba Cloud Model Studio (short audio) ↗qwen-audio-3.1-asr-flash
Alibaba Cloud Model Studio (real-time) ↗qwen-audio-3.1-asr-flash-streaming
QwenCloud (file transcription) ↗qwen-audio-3.1-asr-flash-filetrans

Qwen-Audio (speech) — every version

The full lineage of the Qwen-Audio (speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Qwen-Audio-3.1-TTScurrent2026-09-23—Proprietary (hosted API)
Qwen-Audio-3.1-ASR2026-09-238KProprietary (hosted API)
Qwen-Audio-3.1-TTS-Next2026-09-22—Proprietary (hosted API)

FAQ

What is Qwen-Audio-3.1-ASR?

Qwen-Audio-3.1-ASR is Alibaba's hosted speech-to-text model from the Qwen-Audio-3.1 release of 23 September 2026. It transcribes many languages and Chinese dialects and polishes the output by removing filler words and repetitions.

Which versions are available?

Alibaba lists qwen-audio-3.1-asr-flash for short audio over HTTP, qwen-audio-3.1-asr-flash-filetrans for offline files such as meetings and calls, and qwen-audio-3.1-asr-flash-streaming for real-time recognition.

What does it cost?

The short-audio model lists $0.15 per 1M input tokens and $0.47 per 1M output tokens in Singapore, and $0.113 and $0.382 in China (Beijing). The streaming model lists $0.93 and $0.70 in Singapore. The Qwen launch post describes the ASR price cut as up to 95%.

What does ASR-Next add?

Qwen-Audio-3.1-ASR-Next, released alongside it, labels speakers with timestamps and aligned transcripts, and recognises emotions, ambient sounds and machine noise for sound captioning, event localisation and audio question answering.