← All posts / Models

GLM-5.3: The Open Coding Model That Found 2,436 Real Bugs — and Its Own Weights Delayed

Z.ai's GLM-5.3 gets 50% better at coding from post-training alone, tops the CyberGym vulnerability benchmark at 84.5, and surfaced 2,436 real open-source vulnerabilities — so Z.ai delayed its own open-weights release for safety review.

GLM-5.3: The Open Coding Model That Found 2,436 Real Bugs — and Its Own Weights Delayed

Most AI model releases announce what a model can do. Z.ai’s GLM-5.3, shipped on August 14, 2026, is unusual in that its most newsworthy fact is what the company decided not to do yet: publish the weights. The reason is that the model got dramatically better at something its creators didn’t fully plan for — finding and exploiting real security vulnerabilities — and the lab paused its own open release to run additional safety hardening.

The result is a release that reads like a capability announcement and a safety case study at the same time, and it says a lot about where open-weight AI is heading in 2026.

Same base model, 50% better coding

The first surprise in GLM-5.3 is architectural: there isn’t one. The model reuses the exact same base as GLM-5.2 — a roughly 744-billion-parameter mixture-of-experts with about 40B parameters active per token. There was no new pretraining run. Every improvement came from post-training: reinforcement learning and fine-tuning applied after the fact.

Despite that, Z.ai reports a 50% jump on its in-house Code Bench, and state-of-the-art performance among open-source models on public benchmarks including Terminal-Bench 3.0 (28.3%, up from 4.6% for GLM-5.2) and Agents’ Last Exam (CLI). On the API side, the model supports a 1M-token context window with 128K maximum output, text-only input, and always-on reasoning with three effort levels (low, high, max) — disabling reasoning is no longer supported at all.

A 50% coding gain from post-training on a frozen base is a strong signal of how much headroom current frontier-scale models still have. The capability was latent in GLM-5.2; post-training simply elicited it.

The emergent part: cybersecurity

Then came the part Z.ai itself describes as exceeding expectations. As the scale of post-training expanded, the model’s cybersecurity capabilities improved “at a rate that exceeds expectations.” The deeper along the vulnerability-exploitation chain you measure, the bigger the gains over GLM-5.2 become:

  • CyberGym (UC Berkeley’s vulnerability-discovery benchmark, 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects): GLM-5.3 scored 84.5, ahead of Claude Mythos 5 (83.8) and GPT-5.6 Sol (83.6). An open-weights model at the top of that chart is a first.
  • ExploitBench (Carnegie Mellon + Bugcrowd; 41 patched V8 CVEs, measuring whether a model can carry a bug all the way to a working exploit): GLM-5.3 scored 54.4, versus 24.4 for GLM-5.2 — a 2.2x jump. Closed leaders Claude Mythos 5 (78.0) and GPT-5.6 Sol (76.5) still lead by roughly 24 points.
  • ExploitGym (UC Berkeley RDI + Max Planck Institute; 869 containerized exploitation tasks): GLM-5.3 solved 105 tasks in a 2-hour budget and 130 at 6 hours (15.0%), versus 29/39 (4.5%) for GLM-5.2 — about 3.5x, at roughly half of Claude Mythos 5’s 28.4%.

Two independent data points corroborate the harness: Kimi K3’s 32.2 on ExploitBench lands on top of the 32% measured by the joint UK AI Security Institute / US Center for AI Standards evaluation of July 23, and GLM-5.2’s 24.4 matches their 24%.

One caveat worth knowing: CyberGym has four information levels, and Z.ai’s headline 84.5 is the white-box setting (with ground-truth patch diffs provided), where the frontier cluster saturates — five models finish within 7.3 points of each other. The open-ended discovery level remains far harder. The discriminating results are the exploit benchmarks, and that’s where the real, still-unclosed gap with closed frontier models lives.

2,436 real vulnerabilities, one public ledger

The benchmarks would be a normal release story. The ledger is what made GLM-5.3 headline news. Working with external security teams, Z.ai’s models found:

  • 2,436 vulnerabilities across 269 open-source projects, found since GLM-5.2
  • 1,097 rated medium-to-high severity (107 critical, 990 high)
  • 53 publicly disclosed with assigned CVEs; 2,383 still under coordinated embargo
  • Targets spanning system kernels, operating systems, browser engines and network protocols — including Linux, WebKit, FreeBSD, GStreamer and Suricata
  • The oldest defect dates to 1981, with a mean latency of 26.6 years between a bug being introduced and being found

That 26.6-year average is arguably the single most important number in the release. It’s a direct measurement of how much undiscovered vulnerability is sitting in the software everyone runs: the backlog is generational, and model-driven auditing is now finding it faster than humans ever did.

The disclosure practice also deserves credit — a public ledger at cvd.z.ai with assigned CVEs and coordinated disclosure rather than a press release full of unverifiable counts. That is meaningfully better practice than most capability announcements.

Why the weights are late

Which brings us back to the delay. Z.ai says the model’s offensive-security capability grew faster than intended during training — the model reportedly began reasoning across multiple stages of exploitation, forming coherent plans for complete exploit chains, rather than just finding individual bugs. GLM-5.2 shipped MIT-licensed weights within days of its soft launch; GLM-5.3 is the first release in the series to gate open weights behind extra review, with publication pointed at around the end of August.

Whether or not you take the vendor’s framing at face value, the benchmark data is consistent with it: single-bug discovery (CyberGym) improved modestly, while chained exploitation (ExploitBench, ExploitGym) jumped 2.2–3.5x. The capability moved exactly where the claim says it moved. A lab that trained for bug-finding, measured exploit-chaining, and then paused before publishing weights is behaving consistently with what it says it found.

What it means

For defenders, GLM-5.3 is unambiguously good news: the cost of auditing the enormous body of aging open-source code just dropped again, and the coordinated-disclosure ledger is a model other labs should copy. Independent testers note the model still trails Anthropic’s and OpenAI’s frontier models on the hardest coding benchmarks, and reports say it already flagged a security issue in Cursor, the coding tool recently acquired by SpaceX.

For the open-weights debate, the release is a genuine inflection. The open-weight exploitation gap halved in five weeks — from ~24 to ~24 points behind on ExploitBench is gone; it’s now 24 points on a doubled score. Self-hosting a ~744B MoE at 54.4 ExploitBench, with no vendor able to revoke access mid-audit, becomes a real option once weights land.

And for AI governance, it marks a shift: access-gating of risky capabilities is no longer only a US-lab story. An open-weights lab voluntarily delaying its flagship release over emergent offensive capability makes the “open vs. closed” framing more complicated — the review now happens at the lab, before the weights exist in public at all.

Watch the end of August. When the weights drop, budget-conscious and self-hosting security teams get this generation of capability on their own hardware — and everyone else gets a live test of whether safety-gated open release actually works.