Alibaba (Qwen) · 2026-09-02 · major
Qwen3.8-Max-0902 — Alibaba's 2.4T flagship gets a coding refresh
Qwen3.8-Max-0902 is a new snapshot of Alibaba's 2.4-trillion-parameter flagship, post-trained for coding and agent work. It keeps the 1M-token context, and TechNode reports its front-end CodeArena score rose 22 points to 1,691.

A dated snapshot of Qwen3.8-Max, post-trained for bigger codebases and longer unsupervised agent runs.
Key specs
| Code arena (front end) | 1,691 |
|---|
Quick facts
| Maker | Alibaba (Qwen) |
|---|---|
| Model ID | qwen3.8-max-0902 |
| Alias | qwen3.8-max-2026-09-02 |
| Parameters | 2.4T (MoE) |
| Context window | 1M tokens |
| Max output | 131,072 tokens |
| Input modalities | Text, image, video |
Pricing
| Input | $1.65 / 1M tokens |
|---|---|
| Output | $4.951 / 1M tokens |
| Cache read | $0.137 / 1M tokens |
What is it?
Coding on engineering-scale projects is what changed in Qwen3.8-Max-0902, the snapshot Alibaba published on 2 September 2026 under the model ID qwen3.8-max-0902. The base model is unchanged — the same 2.4-trillion-parameter mixture-of-experts flagship with a 1 million-token context — but Alibaba post-trained it further on coding and Cowork-style collaborative agent tasks.
How does it work?
Alibaba lists three areas of improvement for the 0902 snapshot: longer autonomous development runs on large projects, steadier multi-tool orchestration when an agent has to deliver a task end to end, and sharper native vision for chart reasoning and document parsing. The model takes text, image and video input and returns text, accepting up to 991,808 input tokens and returning up to 131,072. Thinking mode and the full tool ecosystem carry over from Qwen3.8-Max.
Why does it matter?
Dated snapshots let teams pin a model version instead of chasing a moving endpoint, and this one arrives with a measurable jump: TechNode reports the front-end CodeArena score climbed 22 points to 1,691, putting Qwen3.8-Max-0902 first on that leaderboard. Alibaba's listed rate of $1.65 per million input tokens keeps it well under Western frontier pricing, which is the trade agent-heavy teams are weighing.
Who is it for?
teams running long coding and agent jobs
Frequently asked questions
- How much does Qwen3.8-Max-0902 cost?
- Alibaba Cloud Model Studio lists Qwen3.8-Max-0902 at $1.65 per million input tokens and $4.951 per million output tokens in the China (Beijing) region, with cache reads at $0.137 and cache creation at $2.063 per million tokens. Implicit caching is charged at $0.206 per million tokens.
- What changed between Qwen3.8-Max and Qwen3.8-Max-0902?
- Qwen3.8-Max-0902 keeps the same 2.4-trillion-parameter base, 1M-token context, thinking mode and tool support as Qwen3.8-Max, and adds further post-training. Alibaba cites stronger coding on engineering-scale projects, better multi-tool orchestration in agent runs, and more reliable chart reasoning and document parsing from its native vision stack.
- Where can developers call Qwen3.8-Max-0902?
- Alibaba serves Qwen3.8-Max-0902 through Model Studio using the ID qwen3.8-max-0902 or the dated alias qwen3.8-max-2026-09-02. Vercel's AI Gateway also carries it as alibaba/qwen3.8-max-0902, which makes it callable from the AI SDK, Claude Code, Codex, Hermes and plain Chat Completions clients.
- What are the token limits for Qwen3.8-Max-0902?
- Qwen3.8-Max-0902 accepts up to 991,808 input tokens and returns up to 131,072 output tokens inside its 1 million-token context window. In thinking mode the input limit is 983,616 tokens with the same 131,072-token output. Model Studio also caps throughput at 1M tokens per minute and 15,000 requests per minute.
Try it
Call model qwen3.8-max-0902 on Alibaba Cloud Model Studio