← All posts / Models

GLM-5.3 Open Weights Land: Z.ai's Exploit-Hunting Flagship Finally Goes Public

After a two-week safety review, Z.ai's GLM-5.3 open weights are scheduled to hit Hugging Face today — and GLM-5.3-Flash's MIT-licensed weights are already live, turning every self-reported benchmark claim into something anyone can verify.

GLM-5.3 Open Weights Land: Z.ai's Exploit-Hunting Flagship Finally Goes Public

Today is the day a two-week wait ends. Z.ai’s GLM-5.3 — the coding model whose exploit-finding ability got so strong the company delayed its own open-weight release — is scheduled to land on Hugging Face today, August 28, letting independent researchers finally check the startling numbers Z.ai has been reporting about itself. The zai-org/GLM-5.3 page on Hugging Face has listed August 28 as the release date, and Z.ai confirmed on X yesterday that “GLM-5.3’s weights will be released tomorrow.” The flagship weights had not appeared publicly at the time of writing, and Z.ai has cautioned the date could still move — but the companion release, GLM-5.3-Flash, is already live under an MIT license, which means the verification era for this model family has effectively begun.

Why the wait mattered

GLM-5.3 was first released on August 14, 2026, but only through Z.ai’s GLM Coding Plan and ZCode — no API weights, no downloads. The reason was unusual: Z.ai described the delay as its most extensive risk review to date, triggered because the model’s post-training run produced unexpectedly strong cybersecurity capability. According to the company’s own account, post-training produced complete exploit chains that Z.ai “never planned” for, and the model has been credited with finding a serious vulnerability in Cursor, the AI coding editor.

The scale of that capability is documented in Z.ai’s running Security Disclosure Ledger: GLM models have been credited with surfacing 2,436 vulnerabilities across 269 open-source projects since GLM-5.2, with more than 2,300 still moving through coordinated disclosure under embargo. That ledger is itself a striking artifact — a lab publicly tracking the offensive-security side effects of its coding models.

Same skeleton, different animal

The most technically interesting thing about GLM-5.3 is what didn’t change. It reuses the same roughly 750B-parameter mixture-of-experts base model as GLM-5.2 (753B by Hugging Face’s count, with around 40B active per forward pass), meaning every reported capability gain comes purely from scaled-up post-training rather than a larger or retrained base. Z.ai calls it the most capable open-weights coding model it has shipped, with a 50% improvement over GLM-5.2 on the in-house Z.ai Code Bench, where it scores 1,769.

The headline jumps are on agentic and security benchmarks:

  • Terminal-Bench 3.0 (real-world coding and agentic task completion): from 4.6% on GLM-5.2 to 28.3% on GLM-5.3 — a leap large enough that reviewers described it as “a different model wearing the same skeleton.”
  • CyberGym (finding real vulnerabilities in code): 84.5%, which Z.ai reports as edging out Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol.
  • ExploitBench (producing a working exploit, not just identifying a flaw): 54.4%, where it still trails frontier closed models.

Every one of those figures is currently Z.ai’s own reported result. That is precisely why today’s release matters: open weights are what finally let outside researchers check them directly rather than taking the company’s word for it.

Flash is already here — and it was hiding in plain sight

While the flagship’s weights were held back, Z.ai shipped something arguably more consequential for everyday developers. GLM-5.3-Flash, released August 26 with weights already live on Hugging Face (the repository went public August 27 and has drawn well over a thousand likes within a day), is a 320B-parameter MoE that activates only 18B per token — about 5.6% of the network per forward pass, or 8 of 288 experts across its 45 layers.

It is the first natively multimodal model in the GLM-5 series: image and video input, a 1,048,576-token context window, and a newly trained base model built on a 30T-token multimodal corpus. Architecturally, it introduces hybrid attention — combining linear and sparse attention for the first time in the GLM series — and it runs natively in FP8.

Flash also has a fun origin story: it spent its first week running anonymously as “Ox Alpha” on OpenCode and OpenRouter, where developers noticed an unnamed model punching far above its price class — reportedly over 100 tokens/sec at around $0.11 per million input tokens — served entirely on domestically produced Chinese AI chips. The reveal landed earlier this week, and community testing from the anonymous period suggested a 3× improvement in end-to-end serving performance in some setups.

On price, Flash is aggressive even by 2026 standards: list pricing is around $0.15 per million input tokens and $0.50 per million output tokens (OpenRouter has been serving it even cheaper), versus $1.40/$4.40 for the flagship GLM-5.3. Z.ai claims Flash beats GLM-5.2 across benchmarks and real workloads at roughly one-tenth the price, landing within half a point of Claude Opus 4.8 on its internal coding benchmark.

The self-hosting reality check

MIT-licensed weights are one thing; running them is another. The default FP8 Flash checkpoint is roughly 306 GiB of weights before KV cache, and the current vLLM path supports NVIDIA Hopper and newer only. That puts self-hosting in reach of organizations with at least an 8-GPU node (or a GB200 tray at tensor parallelism 4). Everyone below that line consumes Flash as an API — where the economics, not the hardware, are the story.

The flagship is heavier still: expect the ~750B-parameter model to demand a serious multi-node budget, in the same class as GLM-5.2 before it.

Why this matters

Three threads come together in today’s release.

Verification. The industry spent this week arguing about benchmark integrity — silent model swaps behind familiar names at major labs, and benchmark trackers disagreeing about which frontier model is actually smarter. Against that backdrop, an open-weights release of a model with contested, self-reported security scores is a rare chance for the outside world to audit the claims directly.

Post-training as the frontier. GLM-5.3 is evidence that meaningful capability jumps no longer require new base models. If a 6× improvement on Terminal-Bench comes from post-training alone on a year-old skeleton, the base-model arms race is only half the story — and labs that iterate post-training faster get more mileage per training dollar.

Open weights as strategy. Z.ai (formerly Zhipu AI) has now shipped MIT-licensed weights for GLM-5, 5.1, 5.2, and 5.3-Flash, keeping a cadence that keeps Chinese labs within striking distance of the closed frontier — as one analyst put it, “how Chinese labs keep stride with the frontier.” Each release narrows the gap between what’s open and what’s gated behind an API.

What to watch

If the flagship weights appear today as scheduled, watch three things: the license (Flash is MIT, but Z.ai has not confirmed the flagship’s terms), independent CyberGym and ExploitBench replications, and how quickly the community ships quantized builds — Flash GGUF variants appeared within a day of release, and the flagship will get the same treatment. And if the date slips again, that itself will be a signal about what the safety review is finding.

Either way, the burden of proof just shifted. From today, GLM-5.3’s claims are checkable — by anyone with the GPUs to check them.