Alibaba (Qwen) · 2026-08-03 · seismic
Qwen3.8-Max — Alibaba's 2.4T flagship ships official with a full benchmark table
Qwen3.8-Max is Alibaba's new 2.4T-parameter MoE flagship, 95B active. It scores 86.6 on Terminal-Bench 2.1 and 92.6 on GPQA Diamond, prices at $2 in / $6 out per 1M tokens, and open weights land next week alongside a Qwen3.8-27B release.

Qwen3.8-Max lands as Alibaba's official 2.4T-parameter, 95B-active MoE flagship with published benchmarks and $2/$6 per-million-token pricing.
Key specs
| Parameters | 2.4T |
|---|---|
| Active params | 95B |
| Price (input) | $2 / 1M |
| Price (output) | $6 / 1M |
| Context in | 983,616 |
Quick facts
| Maker | Alibaba (Qwen team) |
|---|---|
| Total parameters | 2.4T |
| Active parameters | 95B |
| Context window | 983,616 tokens in / 131,072 out |
| Modalities | Text, image, video, documents |
| API price | $2 in / $6 out per 1M tokens |
| Availability | QwenCloud API (OpenAI + Anthropic wire compat) |
| Open weights | Qwen3.8-Max + Qwen3.8-27B released next week |
Benchmarks
| GPT-5.6 Sol (max) | 88.8% | |
|---|---|---|
| Qwen3.8-Max | 86.6% | |
| Claude Opus 4.8 | 84.6% | |
| Claude Fable 5 | 84.6% |
Pricing
| Input | $2.00 / 1M tokens |
|---|---|
| Output | $6.00 / 1M tokens |
| Cache hit | $0.20 / 1M tokens |
| Cache write | $2.50 / 1M tokens |
What is it?
Qwen3.8-Max is Alibaba's flagship large language model, a 2.4 trillion parameter mixture-of-experts network that activates 95 billion parameters per request. It replaces the July 19 preview that shipped without a benchmark table or context spec, and now runs with a 983,616-token input window, 131,072-token output, and native handling of text, images, video, and documents in a single call. API access is live through Alibaba's QwenCloud today; open weights for Qwen3.8-Max and a smaller Qwen3.8-27B land on Hugging Face and ModelScope next week.
How does it work?
The Max tier uses a sparse mixture-of-experts routing to keep active compute low while total capacity climbs — 95B of the 2.4T parameters fire per token. Alibaba pitches long-horizon agent work as the headline feature, with the release page showing a 10-plus-day autonomous coding trace that went from an empty folder to a production project. Vision is a first-class signal in the planning loop rather than a separate encoder call. The API speaks both the OpenAI and Anthropic wire formats.
Why does it matter?
Qwen3.8-Max's benchmark table puts a top-lab Chinese model in the same coding tier as GPT-5.6 Sol and Claude Opus 4.8 at a fraction of the price — 86.6 on Terminal-Bench 2.1 vs Sol's 88.8 and Opus 4.8's 84.6, and 92.6 on GPQA Diamond. Combined with an imminent open-weight release and a matching 27B open model, this is the first time a fully-priced Qwen flagship posts credible frontier-lab numbers before the weights are public. Alibaba's list price of $2 in / $6 out per million tokens is roughly a third of Fable 5's list rate.
Who is it for?
Coding agents, autonomous-workflow builders, and teams evaluating a frontier alternative to Sol or Opus at Chinese-lab pricing.
Frequently asked questions
- How much does Qwen3.8-Max cost to run through the API?
- Qwen3.8-Max lists at $2 per million input tokens and $6 per million output tokens on Alibaba Model Studio, one flat tier covering the full ~1M-token context. Cache hits bill at $0.20 per million, cache writes at $2.50 per million. Thinking tokens count as output. That is roughly a third of Claude Fable 5's list rate.
- When will Qwen3.8-Max open weights ship?
- Alibaba's release post commits to open weights for Qwen3.8-Max next week on Hugging Face and ModelScope, alongside an open-weight Qwen3.8-27B model for local and small-cluster use. The Max tier will be the largest open-weight Qwen flagship to date. A specific date and the exact license are still to be published.
- How does Qwen3.8-Max compare to GPT-5.6 Sol and Claude Opus 4.8?
- On Alibaba's published table, Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1 against Sol's 88.8 and Opus 4.8's 84.6, and 92.6 on GPQA Diamond. It trails Claude Fable 5 on SWE-bench Pro (67.7 vs 80.0) and FrontierSWE (73.5 vs 88.8). The numbers put Qwen3.8-Max in the same tier as the closed frontier for many coding tasks, at a fraction of the sticker price.
- What can Qwen3.8-Max actually do that the July preview could not?
- The preview shipped without a benchmark table, a disclosed context window, or standard per-token pricing. GA fills all three in: 983,616-token input, 131,072-token output, a full benchmark comparison against Sol/Opus 4.8/Fable 5, and standard API pricing. Alibaba also highlights a 10-day autonomous-coding trace on GitHub as the featured demo.
Try it
QwenCloud API (qwen.ai) — model id qwen3.8-max