Mystery Solved: Ox Alpha Was GLM-5.3-Flash All Along — 320B Open Weights for $0.15/M Tokens
Z.ai unmasked its anonymous stealth model: GLM-5.3-Flash is a 320B-A18B natively multimodal MoE with 1M context and MIT-licensed weights at $0.15/$0.50 per million tokens — and the entire preview ran on Chinese-made AI chips.
One week ago, a model called Ox Alpha appeared on OpenRouter under an anonymous “stealth” provider, benchmarked itself to the top of the coding charts, and set the entire developer community playing detective. On August 26, Z.ai ended the game: Ox Alpha is GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, and the weights are already on Hugging Face under an MIT license.
The reveal converts a week of speculation into a product you can download, self-host, or call over an API today. It also answers the biggest open question from the preview: the model that topped early DeepSWE runs was never a Western lab’s A/B test — it was Z.ai stress-testing a production release on tens of thousands of domestically produced Chinese AI chips, in public, without a name attached.
What GLM-5.3-Flash actually is
The specs read like a deliberately efficient frontier model:
- 320B total parameters, 18B active per token. A mixture-of-experts architecture where each token routes through 8 of 288 experts.
- 1,048,576-token context window. The full 1M tokens, matching the ceiling that defined the Ox Alpha listing.
- Natively multimodal. Text, image, and video input — a first for the GLM-5 series, and the capability that fueled the preview’s screenshot-driven coding demos.
- MIT license, weights on Hugging Face. Not “open-ish.” Not research-only. The same license as MIT-licensed developer tools, commercial use included.
- FP8 native weights plus one MTP draft layer for speculative decoding on the serving side.
The pricing is the part that made the Hacker News thread (871 points and climbing) sit up: $0.15 per million input tokens, $0.03 per million cached input, $0.50 per million output. That is roughly one-tenth of GLM-5.3’s $1.40/$4.40 rate card, and a fraction of what frontier-coding-class models from closed labs charge. On the discounted tier, Z.ai reports $0.045 per task on the Artificial Analysis Intelligence Index v4.1.1, where the model scores 57.
The architecture is where the efficiency comes from
GLM-5.3-Flash starts from a newly trained base model on a 30T-token multimodal corpus, and three design choices do the heavy lifting:
Hybrid attention. For the first time in the GLM series, the 45-layer language model interleaves KDA linear-attention layers with NoPE sparse MLA layers. Linear attention handles local dependency; sparse attention retrieves globally relevant context. The split is what lets a 320B model feel fast at million-token context.
IndexPool. At 1M tokens, retrieval itself becomes the bottleneck — attention over a million keys is where latency goes to die. IndexPool compresses groups of indexer key vectors through weighted pooling to hold down latency and memory. Z.ai reports roughly 3× less attention compute and a 4.4× smaller KV cache versus GLM-5.3.
mHC. Manifold-Constrained Hyper-Connections improve scaling efficiency. Against GLM-4.5 at a similar total parameter count, GLM-5.3-Flash roughly halves both activated parameters and layer count.
The result is a model that lands within half a point of Claude Opus 4.8 on Z.ai’s internal coding benchmark while activating 18B parameters per token.
The benchmarks
Z.ai-reported numbers, with the usual caveat that harnesses differ per test and the model card’s footnotes specify temperature, context limits, and judge models per benchmark:
| Benchmark | GLM-5.3-Flash | Reference |
|---|---|---|
| Terminal-Bench 2.1 | 84.3 | Opus 4.8: 85.0 · GPT-5.6 Terra: 87.4 |
| DeepSWE v1.1 | 63.4 | GLM-5.2: 46.2 |
| AutomationBench | 48.8 | GLM-5.2: 26.2 |
| HLE | 55.3 | — |
| OfficeQA Pro | 62.4 | ahead of Opus 4.8 |
| Z.ai Code Bench v1.0 (max) | 29.0 | Opus 4.8: 29.5 |
The jump over GLM-5.2 is the headline: +17.2 points on DeepSWE v1.1, +22.6 on AutomationBench. Independent measurement from Artificial Analysis broadly corroborates the positioning — 57 on the Intelligence Index, 48.7 output tokens/sec, 1.52s time-to-first-token on Z.ai’s API. Strong intelligence-per-dollar, though reviewers note it trends slow and verbose, and vision remains the weak flank: it trails Gemini 3.7 Flash on BabyVision and MVbench.
The early DeepSWE ~80% figure that launched the Ox Alpha frenzy was, as skeptics suspected, a partial early run. The verified number on the full v1.1 suite is 63.4 — still best-in-class for open weights at this price, but a useful reminder that stealth-preview numbers deserve skepticism in both directions.
The serving story is the underreported part
The most strategically significant detail is buried in the deployment notes: the entire Ox Alpha preview ran on domestically produced Chinese AI chips, using a custom SGLang-based engine that disaggregates encoding, prefill, and decoding, and Z.ai reports a 3× end-to-end serving improvement across tens of thousands of accelerators.
Read that again. A frontier-class model was served at public scale — millions of free tokens to a benchmark-hungry developer community — on hardware that is not NVIDIA’s. Whatever the per-chip performance gap, the aggregate serving capacity exists, it works, and it just survived a week of adversarial community load-testing. That is a data point about export controls and the global compute landscape that no press release could have purchased.
Who can actually self-host
The MIT license is generous; the hardware requirements are not. The default FP8 checkpoint is roughly 306 GiB of weights before KV cache, and the current vLLM path supports NVIDIA Hopper and newer only. Realistic self-hosting means an 8-GPU node (or a GB200 tray at TP4) — mid-size and large organizations, plus AI-native startups renting capacity. Everyone below that line consumes the API, where the economics rather than the hardware are the story.
Local serving is supported on SGLang, vLLM, TokenSpeed, and KTransformers, and unsloth already has a dynamic-quanted variant up. The r/LocalLLaMA megathread filled within hours with DGX Spark owners doing the math on 320B-A18B at FP8 — the model fits the “small cluster, real capability” niche that the community has been begging labs to fill.
On the hosted side, GLM-5.3-Flash is live for all GLM Coding Plan tiers — Lite ($18/mo), Pro ($80), Max ($168) — at 3× the usable quota of GLM-5.3, with multimodal capabilities surfacing in ZCode through Browser Use and Computer Use. It is also live on OpenRouter, closing the loop on where the preview ran.
Why this matters
Three takeaways outlive the reveal itself.
The stealth-drop playbook is now confirmed as a deliberate launch strategy. Z.ai ran a week-long anonymous preview, let the community generate the hype, benchmarked in public for free, and converted the reveal into maximal coverage — exactly the pattern we outlined when Ox Alpha first appeared. Expect every lab with a credible checkpoint to copy it.
Open weights at frontier-adjacent capability keep getting cheaper. A year ago, a 1M-context multimodal model within half a point of Opus on coding benchmarks at $0.15/M input was not a serious product category. GLM-5.3-Flash, DeepSeek’s V4-Flash, and Qwen3.8’s open drops have made it one. The closed-lab price umbrella is eroding from below at a pace that is getting hard to ignore.
The serving detail is a geopolitical signal. Chinese labs are not just training competitive models — they are serving them at scale on domestic silicon, in public, under load. The compute moat everyone assumed may be shallower than advertised.
For developers, the practical move is simple: the free mystery window is over, but the model didn’t get worse — it got an MIT license, a price cut, and a support plan. The canary that was singing anonymously for a week just filed its paperwork.
Sources
- [1] https://z.ai/blog/glm-5.3-flash
- [2] https://huggingface.co/zai-org/GLM-5.3-Flash
- [3] https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/
- [4] https://news.ycombinator.com/item?id=49449507
- [5] https://www.reddit.com/r/LocalLLaMA/comments/1vyzzxu/megathread_glm53flash_former_oxalpha/
- [6] https://openrouter.ai/z-ai/glm-5.3-flash
- [7] https://docs.z.ai/guides/overview/pricing