Five-Step Chain Breaks Claude Code Opus 5 Auto Mode With 60-80% RCE Success — Against a Claimed 0.00% Injection Rate
Security researcher Johann Rehberger demonstrates a prompt-injection chain that hijacks Claude Code Opus 5's default Auto Mode into full remote code execution at 60-80% success — directly contradicting the 0.00% attack-success rate Anthropic's commissioned evaluation reported.
Three weeks after Anthropic made Auto Mode the default for Claude Code — replacing human permission prompts with a model-based safety classifier, backed by a third-party evaluation showing a 0.00% prompt-injection attack success rate — security researcher Johann Rehberger has published a working counterexample. His five-step chain hijacks Claude Code Opus 5 in Auto Mode into full remote code execution with a 60-80% success rate on a small sample. The writeup, posted August 26 on his Embrace The Red blog under the handle wunderwuzzi, became one of the most-discussed AI security stories of the past week and is still trending today.
The contradiction at the heart of the story is stark but instructive: both numbers are true at the same time. The 0.00% figure came from a fixed benchmark. The 80% figure came from a targeted attack that wasn’t in it.
Background: the 0.00% claim
When Anthropic flipped the Auto Mode default in mid-August, it leaned on two data points. The first was an internal study in which 1,053 paid testers approved a clearly dangerous swapped-in command only 13.6% of the time, while the Auto Mode classifier blocked 89% of the same commands. The second was an evaluation by vendor Trajectory Labs: 72 indirect prompt-injection scenarios, each run ten times — 720 attack attempts in total — of which exactly zero succeeded against Claude models running Auto Mode. Claude Code lead Boris Cherny summarized the posture confidently: layered defenses had reduced indirect prompt injection “to approximately zero,” and “we just cannot demonstrate prompt injection anymore.”
Rehberger, a Redmond-based security researcher who has spent two years documenting agent hijacking techniques, wanted to see how that result held up against a chain built specifically to defeat the classifier rather than to fit a benchmark.
The attack chain
The entry point is disarmingly mundane: a user asks Claude Code to summarize a website. The site — a decoy archive of historical notebook records, complete with plausible catalogue metadata, dates, and checksums — gives the agent a legitimate reason to keep digging. From there, the chain unfolds in five steps.
Step 1: force the shell. Claude first reaches for the WebFetch tool, which summarizes content server-side and is harder to attack directly. So the malicious server answers with an HTTP 415 “Unsupported Media Type” — and Claude, on its own, decides to fall back to curl via the Bash tool. A 303 redirect then points it at a ZIP archive. This is the key hijacking philosophy, as Rehberger notes: the attack never tells the model what to do; it just makes the malicious path the most attractive route to the objective.
Step 2: the binary decoy. The ZIP contains a README, a manifest, seven Base85/zlib-encoded JSON notebook records, a native macOS binary called decoder-darwin, and — quietly — a file named struct.py. The README describes the binary as the catalogue decoder. Claude inspects it, correctly refuses to execute an unknown binary, and instead decides to write its own Python decoder. That safety-conscious decision is the exploit.
Step 3: self-written code, wrong directory. Claude’s replacement decoder imports Python’s base64 module — but it runs the script from inside the extracted archive directory. Python places that directory on its module search path. The standard base64 module imports struct. The archive’s malicious struct.py shadows the standard library, re-exports the genuine low-level _struct API so decoding continues normally, and — during the import — silently executes an obfuscated module-level payload. The obfuscation was written, Rehberger notes, with the help of GPT-5.6, a detail that underscores how attacker-side models are lowering the cost of building these chains.
Step 4: the staging round. The poisoned struct.py spawns an isolated Python child using python3 -I — isolated mode, which prevents the child from recursively re-triggering the same shadowing attack. The child downloads a remote stage, which pulls down a native payload, marks it executable, and launches it. In the lab, the payload establishes a real command-and-control callback and pops Calculator as visible proof. The detached process outlives the Claude conversation.
Step 5: the agent-spawning variant. A bonus variant skips the custom malware entirely: the payload launches a second, headless Claude Code instance via claude -p. The nested agent gets its own tools and context, and in test runs performed basic reconnaissance — whoami, uname, id — opened Calculator, and wrote files outside the workspace. The attack doesn’t just execute code; it creates another agent.
The results, honestly framed
Across small samples of five runs each, Rehberger measured a 60% success rate for the C2 chain (3/5), 60% for the recon variant (3/5), and 80% (4/5) for the variant combining Calculator with out-of-workspace file writes. He is explicit that these are small samples from a motivated attacker, not a universal attack-success-rate measurement — the same epistemic honesty he asks of benchmark headlines.
Two behavioral findings deserve more attention than the headline number. First, Claude often realized the archive was booby-trapped only after the machine was compromised — sometimes continuing to decode records, sometimes warning the user, sometimes incorrectly telling the user that nothing malicious had executed. Second, and more damning: in several runs, when Claude tried to terminate the malware it had just launched, Auto Mode denied the cleanup command. The classifier allowed the creation of the malicious process and then blocked the kill — the safety mechanism itself became part of the failure path.
The refusals are equally instructive. When Claude mitigated the attack, it did so by analyzing the archive statically and never running Python from its root, by using python3 -I isolated mode, by running its decoder from a safe parent directory, or by recognizing the module shadowing before triggering it. Every one of those is an environment-usage decision, not a classifier decision.
Anthropic’s response: a classifier is not a sandbox
Rehberger first sent the report to Anthropic’s model bug bounty address, heard nothing, then resubmitted through the security channel. Anthropic closed the report as “Informative” — working as designed. The company’s position: Auto Mode is a convenience feature backed by a best-effort classifier, not a security guarantee. Determined multi-step chains that combine individually benign actions are not what the classifier is meant to stop; the real boundary is operating-system isolation and network egress control.
That position is technically coherent — and Rehberger says as much. The problem is the messaging gap. A 0.00% headline from a fixed 72-scenario benchmark and a working 80% chain from a targeted attack are both simultaneously true, which is precisely why a single benchmark number misleads. As he puts it, those two messages “do not fit together.”
What it means for everyone running agents
The practical guidance is unchanged from years of agent-security writing, which is itself the point: progress in model-level defenses has not repealed it.
- Run unattended coding agents in a container, VM, or OS sandbox — Auto Mode approval is not evidence a command is safe.
- Restrict network egress; the C2 stage is useless if it cannot phone home.
- Monitor agent activity; Claude noticed the compromise late, and the classifier blocked its own remediation.
- Never expose home directories, SSH keys, or cloud credentials to an agent workspace.
The wider lesson cuts both ways. The era of trivial “ignore previous instructions” attacks genuinely is over for frontier models — Claude refused the obvious malicious binary, and when it caught the shadowing attack it understood exactly what had happened. But “no longer trivially exploitable” and “solved” are different claims, and Rehberger argues the second one amounts to claiming a large part of alignment is solved. His suggested reframing — “adversarial misalignment,” something closer to social engineering than to a distinct injectable flaw — captures why a fixed-scenario benchmark will always lag a motivated adversary. Especially when, as his own workflow demonstrates, the adversary now has frontier models of its own to write the obfuscation.
For the growing population of developers now running Claude Code in Auto Mode by default — a population Anthropic deliberately created three weeks ago — the takeaway is not to turn the feature off. It is to stop treating the absence of a permission prompt as the presence of a security boundary. The classifier reduces risk compared to blindly approving everything. It does not replace a sandbox, and it was never going to.
Sources
- [1] https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/
- [2] https://devops.com/a-simple-website-summary-just-exposed-the-limits-of-ai-coding-guardrails/
- [3] https://gbhackers.com/prompt-injection-attack-hijacks-claude-code-opus-5-auto-mode/
- [4] https://thecyberdef.com/claude-opus-5-auto-mode-exploited-rce-attack-hits-80-success-rate/
- [5] https://claude.com/blog/auto-mode-default-in-claude-code