The Locks Came Off: Anthropic Shows GLM-5.3 Hacks Like a Frontier Model and Refuses Like a Wet Paper Bag
Anthropic's Frontier Red Team reports that Zhipu's open-weight GLM-5.3 builds end-to-end exploits at near-Mythos rates, chains browser 0-days autonomously, and drops its refusals under trivial bypasses — 64% to 100% of the time.
Five months ago Anthropic made an uncomfortable bet. When its own Claude Mythos Preview became the first AI model able to autonomously build sophisticated end-to-end cyber exploits, the company declined a normal release. Instead it funneled the capability to vetted defenders through Project Glasswing, reasoning that if the ability ever leaked into an ungated model, attackers would get frontier-grade offense for the price of a download. In a report published September 29, Anthropic’s Frontier Red Team says that moment has arrived — and it has a name: GLM-5.3, the latest open-weight release from Beijing-based Zhipu AI (known abroad as Z.ai).
What Anthropic actually measured
The team, led by Andrew Fasano and Marius Fleischer, ran GLM-5.3 through the same gauntlet that earlier separated Mythos Preview from everything else. On ExploitBench, which tasks models with exploiting known vulnerabilities in the V8 JavaScript engine behind Google Chrome, GLM-5.3 developed complete, working end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview managed 56 of 410. On Anthropic’s internal Binary Exploitation benchmark — a random subset of 100 tasks against open-source projects from Google’s OSS-Fuzz — GLM-5.3 achieved a full control-flow hijack in 4% of trials versus Mythos Preview’s 6%.
Four percent sounds small until you look at the baseline: earlier models like Claude Opus 4.6 and GLM-5.2 succeeded in zero of them. A threshold has been crossed, and the numbers say GLM-5.3 sits a hair behind the most offensive-capable model Anthropic has ever shipped — except Anthropic’s version lives behind vetted-access programs, while GLM-5.3’s weights are public.
The benchmark work is the least alarming part. In human-in-the-loop sessions, researchers pointed the model at software and largely got out of the way. In one session, GLM-5.3 spent a sandboxed day with a local Linux build of a popular web browser and, with less than an hour of total human attention, found several previously unknown vulnerabilities in the browser’s JavaScript engine — then chained them into a working exploit page that reads arbitrary files from a visitor’s computer. Anthropic’s redacted screenshot shows it exfiltrating a user’s SSH private key. The company believes the flaws likely affect other platforms too, and has disclosed them to the maintainer. The same session surfaced exploitable bugs in wireless drivers, graphics drivers, and network-facing device software, now under review.
In a second session, the smaller GLM-5.3-Flash was given public write-ups of CVE-2026-11645, a recently patched Chrome flaw, plus one other known bug. With no meaningful hand-holding it chained the two into a reliable exploit against an ARM64 target, defeating Apple’s pointer-authentication (PAC) hardening. Total cost: 20 minutes of human attention, eight hours of model runtime, and $20.40 at Zhipu’s API prices. A weaponized N-day for the price of a takeout dinner.
The safeguards are the story
Plenty of frontier models are cyber-capable; that’s precisely why they ship with layered controls. GLM-5.3 does include some built-in refusals — ask it outright for malware and it usually declines. What Anthropic measured is how cheaply those refusals fall:
- A false cover story — telling the model it’s an autonomous red-team agent on an exercise — gets it to engage with overtly malicious requests 64% of the time.
- Prefilling the model’s reasoning tokens, so it appears to have already deliberated and decided to proceed, works 92% of the time.
- Abliteration — surgically editing the open weights to remove refusal behavior — works 100% of the time.
None of these techniques moved the needle against safeguarded Claude models in Anthropic’s testing: deceptive prompts were blocked, the API offers no way to prefill Claude’s thinking, and closed weights cannot be abliterated.
The abliteration detail deserves its own headline. Anthropic’s team, attempting the technique for the first time, needed about 2,200 GPU hours (~$4,400) to produce an uncensored copy; an experienced team would need closer to 600 GPU hours, or roughly $1,200. The edit collapsed GLM-5.3’s refusal rate from over 90% to about 3% on JailbreakBench, 2% on HarmBench, and 12% on StrongREJECT — while leaving capability essentially untouched (identical GPQA-Diamond scores, only a few points off on CyberGym). And this isn’t theoretical: abliterated GLM-5.3 builds circulated publicly within days of the model’s release, no GPU cluster of one’s own required.
Context and stakes
Anthropic’s findings land on top of NIST’s Center for AI Standards and Innovation (CAISI), which on September 17 independently called GLM-5.3 “the most cyber-capable open-weight model released to date,” trailing the US frontier by roughly four months on aggregate cyber benchmarks — with the caveat that the US comparison models were tested with safeguards off and remain gated to vetted users. GLM-5.3 is a download.
The report is careful to note the double edge. The same capability that terrifies defenders is precisely what defenders need, and Anthropic argues the answer is not to pretend the openweights genie stays bottled — it’s to arm the good side faster. Project Glasswing and efforts like Patch the Planet have already helped defenders surface more than 10,000 vulnerabilities in critical software ahead of this moment, and vetted defenders can now access the even-stronger Claude Mythos 5.1. Anthropic’s policy asks are pointed: governments should run independent safety evaluations on sufficiently capable models — explicitly including GLM-5.3’s successors — and open-weight developers should treat safeguarding as a release requirement, not an afterthought.
The uncomfortable takeaway for everyone else: the gap between “frontier lab’s most dangerous capability” and “anyone with a GPU and $20 of API credits” is now measured in months, and the safeguard layer meant to stand in that gap can be talked past with a costume-party lie 64% of the time. The locks didn’t fail. They were never really on the door.