█

AI/TLDR

Qwen-Audio-3.1-TTS-Next

Alibaba's audio-creation model from Qwen-Audio-3.1, released 22 September 2026 — speech, sound effects and ambience generated together in one pass.

Qwen-Audio (speech)API only
Released
22 Sep 2026
Parameters
Not disclosed
Input
$0.848 / 1M tokens
License
Proprietary (hosted API)

Overview

Qwen-Audio-3.1-TTS-Next is the audio-creation model Alibaba's Qwen team added in the Qwen-Audio-3.1 release. Alibaba Cloud Model Studio lists its release date as 22 September 2026 and its model id as `qwen-audio-3.1-tts-next`; the Qwen launch post followed on 23 September. It is one of two new models in the release, next to ASR-Next for audio understanding.

Where a standard text-to-speech model reads a script aloud, TTS-Next produces a finished soundtrack. The launch post describes a unified language-model plus diffusion framework that generates voice, sound effects and background audio in one pass, aimed at audiobooks, podcasts, games and ads. The model page lists single-speaker speech, multi-speaker dialogue, podcasts, cinematic soundscapes, ambient audio and sound effects, in Chinese and English.

A request takes up to 3,000 characters of text and up to three reference clips of up to 30 seconds and 10 MB each (WAV, MP3 or OGG Opus) to steer the voices. Output is WAV, MP3 or PCM, up to 120 seconds long, or 240 seconds for podcasts.

The model is served from Alibaba Cloud Model Studio only; no weights are published. In the China (Beijing) region the model page lists $0.848 per 1M input tokens and $1.696 per 1M output tokens, with a limit of 3 requests per second.

Released2026-09-22
LicenseProprietary (hosted API)
WeightsAPI only
ParametersNot disclosed
ModalitiesText, Audio

Pricing

Input$0.848 / 1M tokens
Output$1.696 / 1M tokens

China (Beijing) region rates listed on the model page; 3 requests per second.

Pricing source ↗

Strengths

  • Speech, sound effects and ambient audio generated together from one request
  • Multi-speaker dialogue and podcast output of up to 240 seconds in a single call
  • Up to three 30-second reference clips to steer the voices
  • WAV, MP3 or PCM output, from up to 3,000 characters of input

Best for

  • Reach for it for podcast and audiobook production where narration, dialogue and ambience should come out as one mix
  • Reach for it for game and ad audio that needs sound effects timed with speech
  • Reach for Qwen-Audio-3.1-TTS instead for streaming voice assistants, or for languages beyond Chinese and English

How to access

ProviderModel ID
Alibaba Cloud Model Studio ↗qwen-audio-3.1-tts-next

Qwen-Audio (speech) — every version

The full lineage of the Qwen-Audio (speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Qwen-Audio-3.1-TTScurrent2026-09-23—Proprietary (hosted API)
Qwen-Audio-3.1-ASR2026-09-238KProprietary (hosted API)
Qwen-Audio-3.1-TTS-Next2026-09-22—Proprietary (hosted API)

FAQ

What is Qwen-Audio-3.1-TTS-Next?

Qwen-Audio-3.1-TTS-Next is a hosted audio-creation model from Alibaba's Qwen team, released on 22 September 2026 as part of Qwen-Audio-3.1. It generates speech together with sound effects and background audio in one pass, using a unified language-model plus diffusion framework.

How long can the output be?

The model page lists up to 120 seconds of audio per request, or 240 seconds for podcasts, from up to 3,000 characters of text.

Which languages does it support?

Alibaba Cloud Model Studio lists Chinese and English for Qwen-Audio-3.1-TTS-Next.

What does it cost?

In the China (Beijing) region the model page lists $0.848 per 1M input tokens and $1.696 per 1M output tokens, with a limit of 3 requests per second.

How is it different from Qwen-Audio-3.1-TTS?

Qwen-Audio-3.1-TTS is a streaming speech model for real-time voices in 16 languages. TTS-Next is built for audio production: it mixes speech with sound effects and ambience, supports multi-speaker dialogue and podcasts, and covers Chinese and English.