GLM-5.3: The Open-Weight Model That Got Scary Good at Coding — and Found 2,436 Real Vulnerabilities
Z.ai's GLM-5.3 uses the exact same base model as GLM-5.2 — every gain came from post-training. It jumped from 4.6 to 28.3 on Terminal-Bench 3.0, topped the CyberGym security benchmark at 84.5%, and surfaced 2,436 real vulnerabilities in production code, some 40 years old. Weights land in two weeks.
Z.ai released GLM-5.3 on Friday, August 14, 2026, and the most technically interesting thing about it is what did not change: the base model. GLM-5.3 runs on exactly the same foundation as GLM-5.2. Every improvement — and they are enormous — came from post-training alone. On Terminal-Bench 3.0, the model leapt from 4.6 to 28.3. On DeepSWE v1.1, from 46.2 to 66.9. On Agents’ Last Exam, from 23.8 to 28.5. Same brain, radically better behavior. That is the whole story of where frontier AI progress lives right now: not in bigger pretraining runs, but in what you do after them.
What shipped
GLM-5.3 is Z.ai’s new flagship for complex software engineering and agentic work. The headline specs: a 1M-token context window, 128K maximum output, text-only input, and reasoning that is always on — disabling it is no longer supported. Instead, developers get three reasoning effort levels (low, high, max), and Z.ai recommends max for coding tasks. The model is currently available to GLM Coding Plan subscribers (which starts at $18/month across Lite, Pro, and Max tiers under a new points-based quota system, with off-peak and weekend usage consuming only 50% of standard points). The general API is “coming soon,” and open weights are promised roughly two weeks after launch, following a safety evaluation — a delay the community has grumbled about, since two weeks is a long time in 2026’s AI cycle.
The post-training story
The most revealing part of Z.ai’s technical write-up is how the gains were manufactured. As agent capability improves, the hard part of scaling post-training shifts from the model to the environment. Useful task environments must be executable, verifiable, and close to real professional work — and you need thousands of them, not a handful of hand-built demos.
So Z.ai built pipelines that synthesize environments end to end. Research agents collect task patterns from real work and convert them into runnable long-horizon environments with multi-step dependencies and hidden state. A judge agent then attempts each task to verify it is actually solvable. Verifiers are synthesized without access to the reference solution, and solver trajectories are used to discover and close reward shortcuts. A verifier that passes oracle, no-op, and unsolved-state checks yields a binary reward reliable enough to train on directly. It carries over the RL strategies from GLM-5.2, including SAO with compaction, which helps the gains hold on long-horizon tasks rather than only short ones.
The tasks themselves are not toy coding exercises. In one ML-infrastructure scenario, the model gets the same working environment as a real engineer — compute clusters, storage, internal docs, codebases, experiment results — and must diagnose bottlenecks across the training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness. Some tasks represent several days of work for an experienced engineer. Training on environments at that level pushes the model toward owning substantial work end to end, instead of waiting for users to decompose the problem and supervise each step.
Emergent cyber capability — and a responsible-release dilemma
Here is where GLM-5.3 gets genuinely novel. As part of post-training, Z.ai mixed vulnerability-discovery data and environments into the training set. They expected the model to get better at finding and reasoning about flaws. What surprised them was how quickly the capability kept developing as training scaled: the model began reasoning across multiple stages of exploitation, forming coherent plans for complete exploit chains.
The benchmark picture is a study in contradictions. On CyberGym — white-box source analysis, identifying and validating vulnerabilities by triggering faults — GLM-5.3 scores 84.5%, up from GLM-5.2’s 77.2%. That is the best result on the benchmark, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%), and it makes GLM-5.3 the first open(-ish) model to top a major security benchmark. But on ExploitBench — deeper reasoning about real vulnerabilities and their exploitation — GLM-5.3 reaches only 54.4% (more than doubling GLM-5.2’s 24.4%), while Mythos 5 scores 78.0% and GPT-5.6 Sol 76.5%. On ExploitGym, which measures exploitation tasks completed under time-normalized budgets, GLM-5.3 finishes 105 tasks in two hours and 130 in six, versus 29 and 39 for GLM-5.2 — but Mythos 5 remains well ahead at 181 and 247. Z.ai’s own framing is unusually candid: capability is growing fastest exactly where the model is furthest behind the closed frontier.
Then the benchmarks escaped into the real world. Since GLM-5.2, Z.ai has worked with security teams in China to run its models against production codebases. After expert review, screening, and deduplication, the model identified 2,436 vulnerabilities across 269 open-source projects, including 1,097 rated medium-to-high severity. The findings span system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols — many had gone unnoticed for years or even decades, with the oldest dating back roughly 40 years. VentureBeat also reports the model flagged a potentially serious vulnerability in Cursor, the AI coding editor, during testing. Z.ai has spun this into a standing disclosure program: the Z.ai Security Disclosure Ledger publicly tracks findings as they move through responsible disclosure, recording affected project, severity, CVE where available, and how long the flaw had lurked in the codebase.
This is precisely why the weights are being held back for a safety review. A model this good at finding and chaining exploits, released openly with no gatekeeping, is a dual-use artifact of the first order. Z.ai’s two-week delay is arguably the minimum viable caution.
How it stacks up
On coding, GLM-5.3 claims SOTA among open-source models on public benchmarks, and a 50% gain over GLM-5.2 on the private Z.ai Code Bench. The efficiency numbers matter as much as the scores: at Max effort, GLM-5.3 completes 34.5% of tasks at ~75K output tokens per task, versus GLM-5.2’s 23.4% at 96K. At High effort it hits 31.4% at ~50K tokens — beating Claude Opus 4.8’s 29.5% at 120K. Claude Fable 5 still leads outright at 39.5%. Independent analysis (Interconnects) notes GLM-5.3 surpasses Moonshot’s Kimi K3 on many benchmarks and edges past Claude Fable 5 or GPT-5.6-Sol on some — with the caveat that all cyber scores come from Z.ai’s own testing.
Why it matters
Three takeaways. First, post-training is now the primary axis of frontier competition — GLM-5.3 proves a full model generation of improvement is possible without touching the base model, if you can synthesize enough realistic, verifiable environments. Second, the security capability curve is bending upward faster than anyone predicted, and it cuts both ways: 2,436 real vulnerabilities surfaced is a public good; the same skill set weaponized is a threat. The responsible-release question Z.ai is wrestling with will face every lab this cycle. Third, the open-weight frontier keeps closing the gap with closed labs — and with weights due within two weeks, GLM-5.3 is positioned to become the default self-hosted coding model the moment it drops.
Sources
- [1] https://z.ai/blog/glm-5.3
- [2] https://docs.z.ai/guides/llm/glm-5.3
- [3] https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor
- [4] https://www.interconnects.ai/p/glm-53-how-chinese-labs-keep-stride
- [5] https://kingy.ai/blog/glm-5-3-specs-benchmarks-api-how-to-use/
- [6] https://atoms.dev/blog/glm-5-3-benchmarks-api-coding-open-weights