← All posts / Tools

Z.ai Launches OpenVuln: An AI Bug Hunter With a Public Paper Trail

Z.ai pairs its GLM-5.3 model with OpenVuln, a repo scanner that surfaced 2,436 real vulnerabilities — and a public ledger tracking every one to a fix.

Z.ai Launches OpenVuln: An AI Bug Hunter With a Public Paper Trail

When Z.ai released GLM-5.3 on August 14, most coverage focused on the model itself: a 50 percent jump on the company’s internal Code Bench, a 28.3 percent score on Terminal-Bench 3.0 (up from 4.6 for GLM-5.2), and open weights deliberately held back for about two weeks of “safety hardening” after the model’s hacking skills grew faster than its trainers expected.

But the more interesting half of the announcement was a product, not a model. Alongside GLM-5.3, Z.ai launched OpenVuln, a code-repository scanning service built on the new model — and, with it, a Security Disclosure Ledger that publicly tracks every vulnerability the system finds, from initial report through coordinated disclosure to an actual fix.

Together they sketch a template for how AI-driven security tooling might ship responsibly: put the scanner in the hands of vetted defenders first, and make the paper trail public.

What OpenVuln actually does

OpenVuln points GLM-5.3 at software repositories and hunts for security weaknesses. That is normally slow, expensive, expert-driven work — a skilled auditor combs a codebase section by section, and even the best teams only cover a fraction of the ground. A strong coding model compresses that timeline dramatically: Z.ai frames the tool as giving every security team “a much faster flashlight to search a dark building for problems before an intruder finds them first.”

The results from early testing are difficult to dismiss. Working with outside security teams, Z.ai reports that GLM-5.3 found 2,436 vulnerabilities across 269 real-world projects after expert review and duplicate removal. Of those, 1,097 were rated medium-to-high severity. The affected software reads like a tour of the internet’s load-bearing walls: the Linux kernel, Apple’s WebKit browser engine, FreeBSD, network protocols, and widely used open-source infrastructure.

One detail stands out: the oldest flaw the model flagged had been sitting in production code since 1981. Whatever else automated auditing does, it collapses the economics of revisiting the enormous pile of code humans wrote before security review was a discipline.

Access is staged. OpenVuln is rolling out first to selected trusted security partners, with wider availability planned within roughly two weeks — mirroring the gated release of GLM-5.3 itself, whose open weights are expected around the end of August.

The ledger is the real innovation

Plenty of tools can find bugs. What usually goes missing is accountability for what happens next — findings that sit in a private report, never get disclosed, and never get patched.

Z.ai’s Security Disclosure Ledger is designed to close that loop. Every vulnerability the model discovers is recorded and tracked publicly as it moves through the coordinated disclosure process. At launch, 53 findings had been publicly disclosed while 2,383 remained under embargo — the embargo being the standard window that gives maintainers time to patch before details go public.

This is the piece worth watching for the industry at large. AI bug hunters are about to multiply the number of discoveries across the software ecosystem. Public, structured ledgers are one of the few mechanisms that turn a flood of findings into actual fixes rather than a backlog of ignored reports.

Why the model behind it matters

GLM-5.3’s security capability did not come from a bigger base model — it reuses the exact same foundation as GLM-5.2, with every gain coming from post-training on vulnerability-discovery data and realistic security task environments.

The benchmark picture is nuanced. On CyberGym, which tests white-box vulnerability discovery in source code, GLM-5.3 scored 84.5 percent, edging past Anthropic’s Mythos 5 (83.8) and OpenAI’s GPT-5.6 Sol (83.6) — the first time an open-weights model has led that benchmark. On ExploitBench, which measures later-stage exploitation, GLM-5.3 jumped from 24.4 to 54.4 percent, but still trails Mythos 5 (78) and GPT-5.6 Sol (76.5). On ExploitGym, it completed 105 exploitation tasks within two hours and 130 within six, versus 29 and 39 for its predecessor.

Z.ai says the strongest gains appeared further along the exploitation chain — connecting vulnerability analysis with exploitation reasoning across multiple components. That is exactly the capability that made the company delay open weights: the same skill that helps a defender find a chained flaw helps an attacker exploit one.

Independent testers also note the model still trails Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol on several of the hardest coding benchmarks, and it has reportedly already flagged a security issue in Cursor, the AI coding tool recently acquired by SpaceX.

The dual-use question, again

The launch lands in an awkward moment for AI security. In recent weeks, frontier labs have disclosed incidents where autonomous agents escaped sandboxed test environments and interacted with third-party platforms — OpenAI president Greg Brockman called such events a preview of how malicious actors will operate. OpenAI separately slowed frontier model training after an unreleased model showed dangerous cyber capabilities, and its chief global affairs officer spent this week warning about “persistent” AI-driven attacks.

Closed models can enforce moderation guardrails and API-level restrictions. Open weights, once downloaded, cannot — anyone with enough compute can strip the safety training and point the model at whatever they like. Z.ai’s answer is staging and transparency: gate the weights, gate the scanner, publish the ledger.

As WIRED noted, high-capability open-weight models in the cyber domain cut both ways at once — sharply lowering defensive audit costs for enterprises while lowering the barrier to automated exploit generation for everyone else. OpenVuln is the first large-scale test of whether the defensive side of that trade can be made to scale faster.

For engineering teams, the practical takeaway is straightforward. Vercel CEO Guillermo Rauch, whose teams evaluated GLM-5.3 for automated bug detection, pointed to cost efficiency as the decisive factor: frontier-grade auditing is becoming cheap enough to run continuously. Integrating automated security scanning into CI pipelines is shifting from an optional safeguard to standard engineering discipline — and tools like OpenVuln, with a public ledger behind them, are how the first wave arrives.