← All posts / Models

Z.ai Releases GLM-5.3: Same Base Model, Post-Training That Spawned an Unplanned Cyber Weapon

Z.ai ships GLM-5.3 with every gain coming from scaled post-training on the unchanged 743B GLM-5.2 base — and an emergent exploit-chaining capability the company says it never planned.

Z.ai Releases GLM-5.3: Same Base Model, Post-Training That Spawned an Unplanned Cyber Weapon

On August 14, 2026, Beijing-based Z.ai released GLM-5.3 — and the most interesting thing about the launch is what did not change. The 743-billion-parameter base model is identical to GLM-5.2’s, untouched. Every improvement Z.ai reports comes from one lever: massively scaled post-training on a broader and more diverse set of professional work environments. The result is being billed as the most capable open-weights model for coding ever measured by the company — plus a cybersecurity capability that Z.ai admits grew faster than it anticipated as training scaled.

The weights are not public yet. Z.ai says they will be downloadable “by anyone” in roughly two weeks, after a safety evaluation and hardening period — a delay the company ties explicitly to the unplanned offensive capability. That two-week gap may end up being the most closely watched part of the whole launch.

Same base, new recipe: environment scaling

GLM-5.2, released in June, introduced the training stack Z.ai built on: IndexShare, a long-context technique; SAO, a reinforcement learning method tuned for long-horizon tasks; and slime, an open-source framework for large-scale asynchronous RL. GLM-5.3’s recipe is simply more — more compute spent across more environments built on that same stack.

Crucially, these environments are designed to resemble units of professional work rather than isolated coding exercises. In one example Z.ai describes, the model is dropped into an ML infrastructure engineer’s working environment — compute clusters, internal documentation, codebases, experiment results — and must diagnose bottlenecks, implement optimizations, and deliver a measurable end-to-end speedup. Some tasks, the company says, represent several days of work for an experienced engineer.

To manufacture environments at volume, Z.ai built pipelines where research agents convert task patterns from real work into runnable long-horizon environments, a judge agent verifies each task is actually solvable, and verifiers are synthesized without access to reference solutions. Machine-generated reward signals cover a subset of tasks, with solver trajectories used to close reward shortcuts. The pipelines still require meaningful human-in-the-loop work — a candid admission in an industry that likes to pretend everything is automated.

The numbers: biggest gains on the longest horizons

The reported results follow exactly the pattern the recipe would predict. On Terminal-Bench 3.0, GLM-5.3 jumps from 4.6 to 28.3 against GLM-5.2. On DeepSWE v1.1, it moves from 46.2 to 66.9. On the CLI variant of Agents’ Last Exam, it climbs from 23.8 to 28.5. All figures are vendor-reported, with methodology footnotes covering harness, context length, and sampling settings.

Against competitors, the picture is mixed. On the in-house Z.ai Code Bench, the company reports a 50% improvement over GLM-5.2 and claims GLM-5.3 outscores Claude Opus 4.8 at comparable effort — 31.4% at roughly 50,000 output tokens per task versus Opus 4.8’s 29.5% at 120,000 — while consuming far fewer tokens. Claude Fable 5 still leads at 39.5% at maximum effort. On public suites including Terminal-Bench 3.0 and DeepSWE, GLM-5.3 trails both GPT-5.6 Sol and Fable 5 on harder coding evaluations. Z.ai argues its private Code Bench reduces contamination risk from public test sets — a fair point, but one that also conveniently can’t be independently checked.

Pricing undercuts the field dramatically: reports put GLM-5.3 API access at roughly $0.11 per million input tokens and $0.28 per million output tokens, with a 1M-token context window carried over from GLM-5.2. It is available now through Z.ai’s API and the GLM Coding Plan, already rolled out to all existing coding-plan subscribers, and usable inside Claude Code, Kilo Code, Cline, OpenCode, and similar harnesses.

The cyber result Z.ai says it didn’t plan

The second headline is the one Z.ai itself flags as unexpected. The team introduced vulnerability-discovery data and environments into post-training expecting incremental improvement at finding and reasoning about individual flaws. Instead, capability kept compounding as training scaled, and the model began reasoning across multiple stages of exploitation — forming coherent plans for complete exploitation chains rather than isolated bug-finding.

The benchmark numbers track that claim. On CyberGym, which tests identifying and validating vulnerabilities from white-box source code, GLM-5.3 scores 84.5%, up from GLM-5.2’s 77.2% and ahead of every model in Z.ai’s comparison set. On ExploitBench, which demands deeper reasoning about real vulnerabilities, it more than doubles its predecessor: 54.4% versus 24.4%. On ExploitGym, which counts exploitation tasks completed under time-normalized budgets, GLM-5.3 finishes 105 tasks within two hours and 130 within six — against 29 and 39 for GLM-5.2.

The frontier gap remains: Mythos 5 completes 181 and 247 ExploitGym tasks on the same budgets. But Z.ai’s own summary of the pattern is blunt — the further up the exploitation chain a benchmark sits, the larger the gain over GLM-5.2, and the wider the remaining gap to closed frontier models.

2,436 real vulnerabilities, 1,097 critical

The capability is not confined to benchmarks. Working with several security teams in China, Z.ai says its models have identified 2,436 vulnerabilities across 269 open-source projects since GLM-5.2 — including 1,097 rated critical or high severity — spanning system kernels, operating systems, browser engines, and network protocols. Many had gone unnoticed for years; the oldest, the company says, was introduced in 1981.

Findings feed a public Z.ai Security Disclosure Ledger that tracks each issue through coordinated disclosure: 53 were publicly disclosed with CVEs assigned at launch, while 2,383 remain under embargo. Recent entries include a use-after-free in the Linux kernel, a WebKit memory-handling flaw affecting Apple Safari, and a parameter-validation bug in FreeBSD. It is the constructive counterpart to a week in which OpenAI described its own test models breaching Hugging Face in a red-team exercise: the same skill that chains exploits also surfaces decades-old bugs for patching.

What to watch

Two things will decide how much of this launch is real. First, independent replication: whether third-party evaluators confirm GLM-5.3’s numbers — particularly the in-house Code Bench results and cyber scores run in Z.ai’s own harness configurations — or expose them as evaluation choice. Second, the weight release around the end of August: an MIT-licensed open-weights model with genuine exploit-chaining capability changes the security calculus for everyone, since analysis of GLM-5.2’s open weights already showed near-frontier autonomous attack capability for as little as $46 per full cyberattack simulation.

One breaking API change is worth noting for developers: GLM-5.3 supports three thinking-effort levels (low, high, max) and no longer permits disabling thinking entirely — applications that previously ran with thinking switched off will need adjustment.

For now, GLM-5.3 stands as the strongest evidence yet that post-training scaling on diverse, realistic environments — not just bigger pretraining runs — can deliver frontier-adjacent capability. And as an unplanned side effect, a reminder that capabilities can emerge from training in ways their creators don’t fully anticipate.