█

AI/TLDR

Z.ai · 2026-08-26 · major

GLM-5.3-Flash — Z.ai opens the weights of the model that was Ox Alpha

GLM-5.3-Flash is the open-weights release of the stealth model that topped OpenRouter as Ox Alpha. Z.ai calls it the first natively multimodal GLM-5 model: 320B total parameters, 18B active, MIT licensed, with a 1M-token context.

Hugging Face model card banner for zai-org/GLM-5.3-Flash

The stealth model that topped OpenRouter now has a name, a model card and MIT-licensed weights: GLM-5.3-Flash.

Quick facts

MakerZ.ai (Zhipu)
LicenseMIT
Parameters320B total, 18B active
Context window1M tokens
ModalitiesText and images
Pretraining corpus30T multimodal tokens
API model IDglm-5.3-flash

What is it?

GLM-5.3-Flash is a 320B-parameter mixture-of-experts model that activates only 18B parameters per token, and Z.ai calls it the first natively multimodal model in the GLM-5 series — it reads images as well as text. It is the model that ran on OpenRouter for a week under the anonymous name Ox Alpha, where it reached the top of the usage leaderboard with more than double DeepSeek's traffic. Z.ai published the blog post, the model card and the API documentation on August 26, 2026.

How does it work?

A hybrid attention design mixes sparse and linear attention, which is how Z.ai keeps cost down at a 1M-token context. Manifold-Constrained Hyper-Connections are used to improve scaling efficiency, and pretraining ran on a 30T-token multimodal corpus. Native multimodal visual coding lets the model look at an interface or a rendered result and act on what it sees, rather than working from text descriptions of the screen.

Why does it matter?

MIT-licensed weights mean a team can self-host the same model it was testing for free on OpenRouter, which is the difference between trialling a 1M-token context and building on it. Z.ai says GLM-5.3-Flash beats GLM-5.2 across benchmarks and real workloads at one-tenth the price, and approaches Claude Opus 4.8 on coding and agentic benchmarks. The weights run under SGLang, vLLM, TokenSpeed and KTransformers.

Who is it for?

teams running long-context coding and agent workloads

Frequently asked questions

Is GLM-5.3-Flash the same model as Ox Alpha?
Yes. Ox Alpha was the anonymous name GLM-5.3-Flash ran under on OpenRouter, listed by a provider called Stealth with no lab attached. Z.ai confirmed to Bloomberg News on August 26, 2026 that the model was a new GLM-series release and that the weights would follow, then published GLM-5.3-Flash the same day with a blog post and a Hugging Face model card.
How does GLM-5.3-Flash compare with GLM-5.3?
GLM-5.3, released on August 14, 2026, is text-only and reuses the GLM-5.2 base model. GLM-5.3-Flash is a separate 320B mixture-of-experts model that reads images as well as text, activates 18B parameters per token, and is built on a hybrid sparse and linear attention design. Z.ai's documentation gives GLM-5.3-Flash three times the quota of GLM-5.3.
What hardware or serving stack does GLM-5.3-Flash need?
Z.ai lists four supported local deployment paths for GLM-5.3-Flash: SGLang, vLLM, TokenSpeed and KTransformers, each with its own cookbook or recipe. Only 18B of the 320B parameters are active per token, so serving cost tracks the active count rather than the full weight size, but the whole checkpoint still has to be held or offloaded somewhere.
Can I use GLM-5.3-Flash without downloading the weights?
Yes. GLM-5.3-Flash is served through the Z.ai API platform under the model ID glm-5.3-flash, documented at docs.z.ai, and it also ran on OpenRouter during its free stealth week as Ox Alpha. Z.ai's API guide describes a points-based quota rather than a flat per-token rate, and it grants three times the quota available for GLM-5.3.

Try it

Model: zai-org/GLM-5.3-Flash on Hugging Face; API model ID glm-5.3-flash

Sources · 5 outlets

Tags

  • model
  • glm-5-3-flash
  • z-ai
  • zhipu
  • glm
  • ox-alpha
  • open-weights
  • mit-license
  • multimodal
  • mixture-of-experts
  • sparse-attention
  • linear-attention
  • long-context
  • coding
  • agentic
  • china

← All releases · Learn AI