← All posts / Models

Ox Alpha Unmasked: Z.ai's Stealth Model Is GLM-5.3-Flash, a 320B Open-Weight Multimodal Contender

The anonymous model that topped OpenRouter's usage charts turned out to be Z.ai's GLM-5.3-Flash — a 320B-A18B natively multimodal MoE with 1M context and MIT-licensed weights at a tenth of frontier pricing.

Ox Alpha Unmasked: Z.ai's Stealth Model Is GLM-5.3-Flash, a 320B Open-Weight Multimodal Contender

For about a week, the most talked-about AI model on the internet had no name attached to it. “Ox Alpha” appeared on OpenRouter and OpenCode on August 20, 2026 with no listed creator, free to use, accepting text, images, and video, with a context window of roughly one million tokens. It promptly surged to the top of OpenRouter’s usage leaderboard — the biggest launch in that marketplace’s history — more than doubling DeepSeek’s usage at its peak. Developers ran it through coding agents, benchmark suites, and vibe-check prompts, and a cottage industry of “guess the lab” speculation bloomed across Reddit, X, and Hacker News.

On Wednesday, August 26, the answer landed: Ox Alpha is Z.ai’s GLM-5.3-Flash, the newest iteration of the GLM series from the Beijing-based lab also known as Zhipu. Bloomberg first reported the confirmation, and Z.ai followed up by releasing the model’s open weights on Hugging Face under an MIT license. What began as an anonymous experiment is now a fully documented, downloadable, production-priced model — and one of the more consequential open-weight releases of the summer.

How the unmasking played out

The clues were there from the start. Early users noticed the model occasionally self-identified as GLM, echoing Zhipu’s earlier GLM releases. The timing fit: Z.ai had shipped GLM-5.3, a text-only flagship, on August 14, and the company had openly stated that multimodal capability would arrive in a follow-up. Reddit’s r/LocalLLaMA thread flagged the similarity within days. Nebius co-founder (and noted benchmark-watcher) previously pointed at the GLM-5.3-Flash identification, and Z.ai confirmed to Bloomberg that the codename was inspired by a popular Chinese film recently released there, titled Niu Lai — “Ox Comes.”

The stealth launch itself was a deliberate strategy, reprising a playbook Alibaba and Xiaomi have both used this year: release the model without claiming credit, let it compete purely on merit, and gather unfiltered user feedback before putting your brand’s reputation on the line. Z.ai described Ox Alpha as “a reasoning model designed for coding, sustained agentic work, and production workloads,” suited for “long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.”

It worked. Stripe CEO Patrick Collison — whose company is acquiring OpenRouter — called the stealth release “very impressive.”

What GLM-5.3-Flash actually is

Strip away the mystery marketing and the specs are substantial:

  • Architecture: Mixture-of-Experts with 320B total parameters and 18B active per token, natively multimodal (text, images, video, visual documents, and interleaved multimodal inputs)
  • Context: 1,048,576 tokens — a full 1M-token window
  • Attention: The first open-source frontier model to combine sparse attention and linear attention in one hybrid architecture — 45 layers interleaving KDA linear-attention layers with NoPE sparse MLA layers, routing each token through 8 of 288 experts, with native FP8 weights and one MTP draft layer
  • Training: A newly trained base on a 30-trillion-token multimodal corpus — notably not a reuse of the GLM-5.2 base, unlike GLM-5.3 itself
  • License: MIT, with weights on Hugging Face from day one
  • Efficiency: Z.ai reports roughly 3× less attention compute and a 4.4× smaller KV cache versus GLM-5.3, aided by a component called IndexPool that compresses groups of indexer key vectors to keep million-token latency and memory in check

The multimodality is the structural differentiator within the GLM-5 family. GLM-5.3 is text-focused; GLM-5.3-Flash bakes vision into training rather than bolting it on. Z.ai trained it to inspect rendered interfaces, gameplay, and 3D output, then assess and revise its own work from that visual feedback — the “look at the result, then fix it” loop that agentic coding actually needs. The same capability extends to documents, spreadsheets, and presentations, with the model able to deliver finished PPTX, PDF, DOCX, and XLSX output end to end.

The numbers

Z.ai reports that GLM-5.3-Flash beats GLM-5.2 across coding and agent suites at roughly one-tenth the price — and the gains are large rather than incremental:

BenchmarkGLM-5.3-FlashReference
Terminal-Bench 2.184.3Claude Opus 4.8: 85.0 · GPT-5.6 Terra: 87.4
DeepSWE v1.163.4GLM-5.2: 46.2
AutomationBench48.8GLM-5.2: 26.2
HLE55.3—
OfficeQA Pro62.4ahead of Opus 4.8
Z.ai Code Bench v1.0 (max effort)29.0Opus 4.8: 29.5

On the independent Artificial Analysis Intelligence Index v4.1.1 it scores 57 — at $0.045 per task on the discounted tier, which Z.ai says pushes the Pareto frontier of intelligence-per-dollar. Independent testing also measured 48.7 output tokens/sec and 1.52s time-to-first-token on Z.ai’s API.

The honest caveats: these are largely vendor-reported numbers with differing harnesses, internal benchmarks like Code Bench can’t be independently replicated, and the model doesn’t claim to top GPT-5.6 Sol or Anthropic’s Fable 5 on the hardest public suites. Vision is its weak flank — it trails Gemini 3.7 Flash on BabyVision and MVbench. This is a value play and an architecture showcase, not a clean leaderboard sweep.

Pricing that changes the math

Standard API pricing is $0.15 per million input tokens, $0.03 per million cached input, and $0.50 per million output — with some third-party providers (OpenRouter) listing it even lower at $0.075/$0.25. Compare that to frontier-model rates and the pitch becomes obvious: you can run long-horizon agent loops all day without watching a token meter. LLM-Stats’ comparison against Claude Sonnet 5 puts GLM-5.3-Flash at roughly 16.8× cheaper on a blended 3:1 basis.

The model is live for all GLM Coding Plan tiers — Lite ($18/mo), Pro ($80), Max ($168) — with 3× the usable quota of GLM-5.3, and its multimodal capabilities surface in ZCode through Browser Use and Computer Use.

The underreported story: Chinese chips

One detail deserves more attention than it has gotten: Z.ai states the entire Ox Alpha preview ran on domestically produced Chinese AI chips. The company used a custom SGLang-based serving engine that disaggregates encoding, prefill, and decoding, and reports a 3× end-to-end serving improvement across tens of thousands of accelerators.

That matters beyond one model. If a top-tier open-weight model can be served at scale on domestic silicon — invisible to users, indistinguishable in quality — then export controls are constraining the supply of chips without constraining the supply of capable models. It also explains some of the pricing aggression: the cost curve looks different when you are not bidding against every frontier lab for the same H-class GPUs.

Self-hosting reality check

The weights are open, but deployment is not trivial. The default FP8 checkpoint is roughly 306 GiB before KV cache, and the current vLLM path supports NVIDIA Hopper and newer only. Realistic self-hosting means a mid-size or large organization with at least an 8-GPU node (or a GB200 tray at TP4), or an AI-native startup renting capacity. Everyone below that line consumes it as an API — where the economics, not the hardware, are the story. Local serving is supported on SGLang, vLLM, TokenSpeed, and KTransformers, and it’s already on Ollama’s library.

Why it matters

Ox Alpha’s week of anonymity ended with a point proven: an unbranded model from a Chinese lab could show up with no reputation, no marketing, and no price tag — and win the usage charts on merit alone. Now unmasked as GLM-5.3-Flash, it joins Kimi K3, DeepSeek’s lineup, and Alibaba’s Qwen series in a summer-long pattern: benchmark-competitive open-weight models at prices Western frontier labs can’t approach, shipped faster, and increasingly trained and served on domestic infrastructure.

For developers, the practical takeaway is simple. If you build coding agents, terminal agents, computer-use workflows, or anything that processes long documents with visual structure, there is now an MIT-licensed, 1M-context, multimodal option that costs about a tenth of what you’re likely paying. For the industry, the question the stealth launch implicitly asked — would you use it if you didn’t know who made it? — got its answer. Enough people said yes that Zhipu’s shares jumped on the reveal.

The next test comes Monday, when Z.ai issues its first detailed earnings report since its January public listing. But the harder question is for OpenAI and Anthropic: when the open tier keeps closing the gap at a tenth of the price, what exactly does the premium buy?