AI/TLDR

Alibaba (Qwen) · 2026-08-03 · seismic

Qwen3.8-Max — Alibaba's 2.4T flagship ships official with a full benchmark table

Qwen3.8-Max is Alibaba's new 2.4T-parameter MoE flagship, 95B active. It scores 86.6 on Terminal-Bench 2.1 and 92.6 on GPQA Diamond, prices at $2 in / $6 out per 1M tokens, and open weights land next week alongside a Qwen3.8-27B release.

Qwen3.8-Max release banner with the 2.4T flagship model announcement
Tech-Now Qwen 3.8 Max review

Qwen3.8-Max lands as Alibaba's official 2.4T-parameter, 95B-active MoE flagship with published benchmarks and $2/$6 per-million-token pricing.

Key specs

Parameters2.4T
Active params95B
Price (input)$2 / 1M
Price (output)$6 / 1M
Context in983,616

Quick facts

MakerAlibaba (Qwen team)
Total parameters2.4T
Active parameters95B
Context window983,616 tokens in / 131,072 out
ModalitiesText, image, video, documents
API price$2 in / $6 out per 1M tokens
AvailabilityQwenCloud API (OpenAI + Anthropic wire compat)
Open weightsQwen3.8-Max + Qwen3.8-27B released next week

Benchmarks

Terminal-Bench 2.1
GPT-5.6 Sol (max)88.8%
Qwen3.8-Max86.6%
Claude Opus 4.884.6%
Claude Fable 584.6%
source ↗
SWE-bench Pro
Claude Fable 580%
Qwen3.8-Max67.7%
source ↗
FrontierSWE
Claude Fable 588.8%
Qwen3.8-Max73.5%
Qwen3.7-Max40.7%
source ↗
GPQA Diamond
Qwen3.8-Max92.6%
Qwen3.7-Max92.4%
source ↗

Pricing

Input$2.00 / 1M tokens
Output$6.00 / 1M tokens
Cache hit$0.20 / 1M tokens
Cache write$2.50 / 1M tokens
source ↗

What is it?

Qwen3.8-Max is Alibaba's flagship large language model, a 2.4 trillion parameter mixture-of-experts network that activates 95 billion parameters per request. It replaces the July 19 preview that shipped without a benchmark table or context spec, and now runs with a 983,616-token input window, 131,072-token output, and native handling of text, images, video, and documents in a single call. API access is live through Alibaba's QwenCloud today; open weights for Qwen3.8-Max and a smaller Qwen3.8-27B land on Hugging Face and ModelScope next week.

How does it work?

The Max tier uses a sparse mixture-of-experts routing to keep active compute low while total capacity climbs — 95B of the 2.4T parameters fire per token. Alibaba pitches long-horizon agent work as the headline feature, with the release page showing a 10-plus-day autonomous coding trace that went from an empty folder to a production project. Vision is a first-class signal in the planning loop rather than a separate encoder call. The API speaks both the OpenAI and Anthropic wire formats.

Why does it matter?

Qwen3.8-Max's benchmark table puts a top-lab Chinese model in the same coding tier as GPT-5.6 Sol and Claude Opus 4.8 at a fraction of the price — 86.6 on Terminal-Bench 2.1 vs Sol's 88.8 and Opus 4.8's 84.6, and 92.6 on GPQA Diamond. Combined with an imminent open-weight release and a matching 27B open model, this is the first time a fully-priced Qwen flagship posts credible frontier-lab numbers before the weights are public. Alibaba's list price of $2 in / $6 out per million tokens is roughly a third of Fable 5's list rate.

Who is it for?

Coding agents, autonomous-workflow builders, and teams evaluating a frontier alternative to Sol or Opus at Chinese-lab pricing.

Frequently asked questions

How much does Qwen3.8-Max cost to run through the API?
Qwen3.8-Max lists at $2 per million input tokens and $6 per million output tokens on Alibaba Model Studio, one flat tier covering the full ~1M-token context. Cache hits bill at $0.20 per million, cache writes at $2.50 per million. Thinking tokens count as output. That is roughly a third of Claude Fable 5's list rate.
When will Qwen3.8-Max open weights ship?
Alibaba's release post commits to open weights for Qwen3.8-Max next week on Hugging Face and ModelScope, alongside an open-weight Qwen3.8-27B model for local and small-cluster use. The Max tier will be the largest open-weight Qwen flagship to date. A specific date and the exact license are still to be published.
How does Qwen3.8-Max compare to GPT-5.6 Sol and Claude Opus 4.8?
On Alibaba's published table, Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1 against Sol's 88.8 and Opus 4.8's 84.6, and 92.6 on GPQA Diamond. It trails Claude Fable 5 on SWE-bench Pro (67.7 vs 80.0) and FrontierSWE (73.5 vs 88.8). The numbers put Qwen3.8-Max in the same tier as the closed frontier for many coding tasks, at a fraction of the sticker price.
What can Qwen3.8-Max actually do that the July preview could not?
The preview shipped without a benchmark table, a disclosed context window, or standard per-token pricing. GA fills all three in: 983,616-token input, 131,072-token output, a full benchmark comparison against Sol/Opus 4.8/Fable 5, and standard API pricing. Alibaba also highlights a 10-day autonomous-coding trace on GitHub as the featured demo.

Try it

QwenCloud API (qwen.ai) — model id qwen3.8-max

Sources · 5 outlets

Tags

  • model
  • qwen
  • qwen-3-8
  • qwen-3-8-max
  • alibaba
  • chinese-lab
  • flagship-model
  • moe
  • multimodal
  • coding-model
  • open-weights
  • terminal-bench
  • swe-bench

← All releases · Learn AI