Overview
Qwen-Audio-3.1-Realtime is the live voice-conversation model in Qwen-Audio-3.1, Alibaba's speech release announced by the Qwen team on 23 September 2026. The QwenCloud model changelog lists it on 20 September 2026 under the model id `qwen-audio-3.1-realtime-plus`, describing real-time duplex speech conversations that use the same integration protocol as the 3.0 Plus model, so existing integrations carry over.
It is an end-to-end speech-to-speech model: one connection takes the user's audio and returns spoken and text replies, without separate recognition and synthesis calls. Alibaba Cloud Model Studio describes it as a real-time duplex speech model with text and audio input and output, function calling, web search and voice cloning; web search and function calling cannot be enabled at the same time. It is served over WebSocket, AOQ and WebRTC, with server VAD, smart turn detection or client-controlled turns.
The model has a 262,144-token context window, with up to 245,760 input and 16,384 output tokens, and keeps up to 50 audio turns or 300 seconds of audio as conversation history. It supports German, English, Spanish, French, Indonesian, Italian, Japanese, Korean, Portuguese, Russian and Chinese, plus 20 Chinese regional varieties. The release keeps the earlier voices and adds eight system voices, for 13 in total, alongside cloned voices.
In Singapore the model page lists $0.8 per 1M text input tokens, $6.4 per 1M audio input tokens, $6.4 per 1M text output tokens and $24 per 1M audio output tokens; China (Beijing) lists $0.688, $5.501, $5.501 and $20.628. Both regions allow 60 requests and 100,000 tokens per minute.
| Released | 2026-09-20 |
|---|---|
| License | Proprietary (hosted API) |
| Weights | API only |
| Parameters | Not disclosed |
| Context | 262K |
| Max output | 16K |
| Modalities | Audio, Text |
Pricing
| Input | $6.4 / 1M tokens |
|---|---|
| Output | $24 / 1M tokens |
Singapore audio rates for qwen-audio-3.1-realtime-plus; text is $0.8 in / $6.4 out. China (Beijing) is $5.501 / $20.628 for audio and $0.688 / $5.501 for text.
Strengths
- Full-duplex speech-to-speech in one connection, with no separate ASR and TTS calls
- Function calling, web search and voice cloning in a live voice session
- 262,144-token context window with up to 50 audio turns of history
- 11 languages plus 20 Chinese regional varieties, and 13 system voices
- Same integration protocol as the 3.0 Plus realtime model
Best for
- Reach for it for voice assistants and customer-service agents that must answer and act in the same spoken turn
- Reach for it for browser and mobile voice apps, using the WebRTC or AOQ transport instead of WebSocket
- Reach for it when upgrading an app built on the 3.0 Plus realtime model, since the protocol is unchanged
- Reach for Qwen-Audio-3.1-ASR and Qwen-Audio-3.1-TTS instead when you need transcription or synthesis on their own
How to access
| Provider | Model ID |
|---|---|
| Alibaba Cloud Model Studio ↗ | qwen-audio-3.1-realtime-plus |
| QwenCloud ↗ | qwen-audio-3.1-realtime-plus |
Qwen-Audio (speech) — every version
The full lineage of the Qwen-Audio (speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Qwen-Audio-3.1-TTScurrent | 2026-09-23 | — | Proprietary (hosted API) |
| Qwen-Audio-3.1-ASR | 2026-09-23 | 8K | Proprietary (hosted API) |
| Qwen-Audio-3.1-TTS-Next | 2026-09-22 | — | Proprietary (hosted API) |
| Qwen-Audio-3.1-Realtime | 2026-09-20 | 262K | Proprietary (hosted API) |
FAQ
What is Qwen-Audio-3.1-Realtime?
Qwen-Audio-3.1-Realtime is Alibaba's hosted full-duplex speech-to-speech model from the Qwen-Audio-3.1 release. It holds live spoken conversations with audio and text input and output, and supports function calling, web search and voice cloning. Its model id is qwen-audio-3.1-realtime-plus.
When was it released?
The QwenCloud model changelog lists qwen-audio-3.1-realtime-plus on 20 September 2026. The Qwen team announced the full Qwen-Audio-3.1 release on 23 September 2026.
Which languages does it support?
The developer guide lists German, English, Spanish, French, Indonesian, Italian, Japanese, Korean, Portuguese, Russian and Chinese, plus 20 Chinese regional varieties.
What does it cost?
In Singapore the model page lists $0.8 per 1M text input tokens, $6.4 per 1M audio input tokens, $6.4 per 1M text output tokens and $24 per 1M audio output tokens. China (Beijing) lists $0.688, $5.501, $5.501 and $20.628, with 60 requests and 100,000 tokens per minute in both regions.
Do I need to change code to move from the 3.0 realtime model?
The changelog says Qwen-Audio-3.1-Realtime-Plus uses the same integration protocol as 3.0 Plus, keeps the existing voices and adds eight system voices, so switching is mainly a change of model id.