GLM-5.3's Hacking Skills Outgrew Z.ai's Expectations — and Open Weights Are on Hold
Z.ai's new coding model found 2,436 real vulnerabilities in production software and became the first GLM release to hold back open weights for safety review.
When Z.ai released GLM-5.3 on August 14, 2026, the headline was supposed to be coding: a 50% jump on the company’s in-house Code Bench, achieved with the exact same base model as GLM-5.2 — every improvement came from post-training. But the story that has kept the release in the news for the past week is something the company admits it did not fully plan for. The model’s cybersecurity capabilities, seeded through vulnerability-discovery data in the training mix, kept growing as training scaled — far beyond what Z.ai expected — and the model began assembling complete exploitation chains on its own. The consequence is a first for the GLM series: open weights, usually shipped alongside a release, are being held back for roughly two weeks of safety review and hardening.
A 50% coding jump from post-training alone
GLM-5.3 shares its base model with GLM-5.2. All of its gains come from what happened after pre-training, which makes the scale of the improvement notable. On Z.ai’s internal Code Bench, the model scores 50% higher than its predecessor. On public benchmarks, Z.ai reports state-of-the-art results among open-source models on Terminal-Bench 3.0 — where it scores 28.3% against GLM-5.2’s 4.6% — and on Agents’ Last Exam (CLI).
The technical details matter for developers evaluating a switch. GLM-5.3 is text-only, with a 1M-token context window and 128K maximum output. Reasoning is always on and can no longer be disabled; instead, three effort levels (low, high, max) modulate depth, with max recommended for complex coding work. The API is served through OpenAI Chat Completion, OpenAI Response, and Anthropic Message-compatible protocols, and the GLM Coding Plan — starting at $18 per month — uses a points-based quota where off-peak calls, including all weekend usage, consume half the standard points.
Z.ai attributes the coding gains to a shift in how training environments were built: tasks that resemble “real units of expert work” rather than coding exercises, some representing several days of work for an experienced engineer — diagnosing bottlenecks across a training stack, running experiments, and delivering a measurable end-to-end speedup while preserving correctness.
The surprise: emergent exploitation capability
The unexpected part sits in the security domain. As part of post-training, Z.ai introduced vulnerability discovery data and environments, expecting a better bug-finder. What they observed was qualitatively different. GLM-5.3 did not just get better at identifying isolated flaws — it began reasoning across multiple stages of exploitation, forming coherent plans for complete exploitation chains.
Three benchmarks map the picture. On CyberGym, which starts from white-box source code and tests whether a model can identify and validate vulnerabilities by triggering faults, GLM-5.3 scores 84.5% — up from GLM-5.2’s 77.2% and the best result on the benchmark, ahead of Anthropic’s Mythos 5 at 83.8% and OpenAI’s GPT-5.6 Sol at 83.6%. On ExploitBench, which demands deeper reasoning about real vulnerabilities and their exploitation, GLM-5.3 more than doubles its predecessor to 54.4% — but still trails Mythos 5 (78.0%) and GPT-5.6 Sol (76.5%) by roughly 24 points. On ExploitGym, which counts completed exploitation tasks under time-normalized budgets, GLM-5.3 finishes 105 tasks in two hours and 130 in six, against 29 and 39 for GLM-5.2 — with Mythos 5 still well ahead at 181 and 247.
Z.ai’s own summary of the pattern is unusually candid: the further up the exploitation chain a benchmark sits, the larger the gain over GLM-5.2 — and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where the open model is furthest behind.
2,436 real vulnerabilities, one 40 years old
Benchmarks are one thing; production software is another. Since GLM-5.2, Z.ai has been working with security teams in China to run its models against real-world codebases. After expert review, screening, and deduplication, GLM-5.3’s testing program identified 2,436 vulnerabilities across 269 projects, including 1,097 rated medium-to-high severity. The findings span system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols — software that includes the Linux kernel, the WebKit browser engine, and FreeBSD.
Perhaps the most striking detail: many of these bugs had sat unnoticed for years or even decades, with the oldest dating back roughly 40 years — a flaw introduced into code around 1981, before some of the people now fixing it were born. VentureBeat additionally reported that the model flagged a serious vulnerability in Cursor, the AI coding tool, during testing.
That discovery work has grown into the Z.ai Security Disclosure Ledger, a public record that tracks findings through the disclosure process, distinguishing already-public issues from those still under coordinated disclosure. For disclosed vulnerabilities, the ledger records the affected project, severity, CVE where available, and how long the flaw had lived in the codebase.
Why the weights are being held back
For a lab whose brand is built on open weights, delaying a flagship release is a significant choice. Z.ai says the model’s exploitation skill grew faster than anticipated during training, and that holding weights for about two weeks of additional safety review is a deliberate gate — the first time the GLM series has done this. The public weights are now pointed at around the end of August, which is when self-hosting teams will be able to run this generation on their own hardware. Until then, access runs through the GLM Coding Plan and Z.ai’s ZCode tool.
Alongside the model, Z.ai launched OpenVuln, a repository scanner built on GLM-5.3. Rather than opening it to everyone at once, the company is rolling it out to selected trusted security partners first, with wider access planned within two weeks — the same staged-release logic applied to the model itself. Z.ai frames OpenVuln as defensive: a much faster flashlight for security teams searching a dark building for problems before an intruder finds them. The public disclosure ledger, tracking every finding through to a fix, is meant to ensure the work doesn’t just generate a list of exploitable weaknesses that sits unaddressed.
The trade-off nobody has solved
The episode crystallizes a dilemma the whole industry is now confronting in public. A model that can find and reason about vulnerabilities is valuable in direct proportion to its ability to exploit them — the same skill that patches a bug can weaponize it. OpenAI launched Codex Security, an agent for finding and fixing code vulnerabilities, in the same season; Anthropic’s models have demonstrated comparable dual-use behavior in published evaluations. What makes GLM-5.3 distinctive is the transparency of the failure mode: Z.ai expected a better bug-finder and got a system that plans multi-stage exploits, and it responded by slowing the release of the open weights that define its reputation.
Independent testers note that GLM-5.3 still trails Anthropic’s and OpenAI’s frontier models on the hardest coding benchmarks. But the CyberGym result — where an open-weights-lineage model tops the leaderboard, ahead of Mythos 5 and GPT-5.6 Sol — shows the gap is not uniform. And the trajectory is the part worth watching: each post-training cycle is now compounding security capability faster than the headline coding numbers suggest.
For defenders, the calculus is mostly good news: 1,097 medium-to-high severity findings in widely deployed infrastructure, moving through coordinated disclosure with a public ledger, is a genuine win for the ecosystem. For policy on open weights, it is a harder case — the argument for delaying weights used to be theoretical, and now it has a benchmark table attached. Either way, the end of August, when the hardened weights are due, will be one of the more closely watched open-source releases of the year.
Sources
- [1] https://z.ai/blog/glm-5.3
- [2] https://docs.z.ai/guides/llm/glm-5.3
- [3] https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor
- [4] https://www.techtimes.com/articles/324426/20260814/glm-53-post-training-produced-exploit-chains-zai-never-planned-finds-1097-critical-bugs.htm
- [5] https://unrot.co/blogs/today-top-ai-news-august-23-2026