← All posts / Meta

Beltdown: One Untrusted Repo, No Permission Prompt — The Sandbox Escape Anthropic Took 50 Days to Fix

Stealth startup Accomplish disclosed a chain of sandbox holes in Claude Code, OpenAI Codex and Cursor. The Claude Code 'Beltdown' escape ran attacker commands outside the macOS sandbox with zero prompts — and sat unpatched for 50 days and ~30 releases.

Beltdown: One Untrusted Repo, No Permission Prompt — The Sandbox Escape Anthropic Took 50 Days to Fix

If you use an AI coding agent, the sandbox that’s supposed to contain it is the single security boundary that matters. This week, a stealth startup called Accomplish pulled back the curtain on just how leaky those boundaries still are across the three most popular coding agents — Claude Code, OpenAI Codex, and Cursor — and how unevenly the big labs respond when researchers hand them working escapes.

The disclosure, reported exclusively by Alex Konrad at Upstarts Media on September 10 and detailed in technical writeups on Accomplish’s own blog, lands at an awkward moment. The industry is digesting news that a U.S. Senate subcommittee is reportedly investigating OpenAI’s agents breaching Hugging Face systems, and frontier models “escaping” sandboxes to cheat on benchmarks has become a running story. Accomplish’s founders aren’t worried about the headline escapes, though. They’re worried about the long tail of normalized, quietly-patched holes in the tools millions of developers run every day.

Who found this, and why it matters

Accomplish is a Tel Aviv-based startup founded by Or Hiltch (CTO), Amit Avner (CEO), and Guy Zipori, still in stealth, building a hardened way to run AI agents on enterprise endpoints. Their principal security researcher, Oren Yomtov, spent the summer flagging sandbox vulnerabilities to Anthropic, OpenAI, and Cursor. Their position is blunt: new AI models have made formerly low-priority vulnerabilities high-risk, because there is now an autonomous, probabilistic process sitting inside the sandbox that untrusted content can steer.

“There’s a lot of talk about security now,” Hiltch told Upstarts. “It doesn’t really reflect in how they actually build products.”

The response-time gap is the story’s sharpest edge. A vulnerability flagged to Cursor in July, and two reported to OpenAI, were each fixed in about a week. A similar vulnerability flagged to Anthropic took roughly 50 days and about 30 software releases to fully patch. During that window, in Avner’s words, hackers, cyber criminals, or state actors could have been exploiting the gap. OpenAI provided a statement thanking the researchers and noting both issues were addressed in August; Anthropic and Cursor did not respond on the record.

Beltdown: the anatomy of the Claude Code escape

The most striking of the disclosed bugs — dubbed Beltdown — targets Claude Code’s macOS sandbox, which wraps the agent’s Bash tool in Apple’s Seatbelt profile. The counterintuitive twist: turning the sandbox on is what silences the permission prompts. Once Claude sandboxes its commands, it stops asking before running them.

Accomplish’s demo is chilling in its simplicity: they enabled the sandbox, set permission mode to the strictest “don’t ask” setting, opened an untrusted repository in Claude Code, and sent one short message. A command planted in that repo ran on the Mac outside the sandbox, as the privileged user, with no permission prompt at all. It also works via indirect prompt injection — meaning a malicious webpage or document the agent reads could trigger it.

The attack chain exploits a structural gap: Claude Code’s harness runs its own git commands outside the sandbox, in the background, to index the repository. Git’s core.fsmonitor config setting is executed as a shell command whenever git looks at the working tree. So the game becomes: plant a poisoned core.fsmonitor in a .git/config, then get the harness to run a git command that reads it.

Accomplish documented four defenses that stood in the way — and four bypasses:

  1. Hardened git commands. The git calls Claude Code runs outside the sandbox pass flags that blank out core.fsmonitor. But one call was left unhardened: git ls-files, the command that refreshes the file index.
  2. Protected .git folder. Claude’s file tools refuse to write inside .git, Seatbelt blocks bash from writing there, and renaming a folder into .git is blocked — but only for the project root. The Seatbelt rule blocking renames of nested .git folders was missing, so a setup script can build the poisoned git folder under a different name in a subfolder and rename it.
  3. Git runs in the project root, where .git is clean — except the harness actually runs git in whatever directory the Bash tool used last. The script simply changes the working directory into the subfolder.
  4. Nothing re-runs the index refresh on demand — but skill auto-loading does. The script’s last line tells Claude to read a “build report” in that subfolder; when Claude reads a file it checks the folder for skills, and loading one auto-triggers the unhardened git ls-files refresh. The moment Claude reads the report, the payload fires. No user interaction anywhere.

The disclosure timeline tells the patch-race story: reported to Anthropic on July 13, triaged the same day; a first hardening attempt shipped August 6 in v2.1.223 but missed some git calls, so the escape simply moved to another one; the full fix landed August 26 in v2.1.247. Today, every git command the harness runs blanks core.fsmonitor.

SharedRoot: the earlier Cowork escape and the design argument

Beltdown wasn’t Accomplish’s first rodeo. In July they published SharedRoot, an escape from Anthropic’s Claude Cowork agent environment that chained a Linux kernel bug (CVE-2026-46331, “pedit COW” in the act_pedit module) into full host compromise: from a connected folder, the agent reached guest-root via an unprivileged user namespace, then wrote files anywhere on the host Mac — SSH keys, cloud credentials — because Cowork’s local mode mounted the entire host filesystem read-write into the VM at /mnt/.virtiofs-root.

Anthropic closed that report as “Informative,” and Cowork now uses cloud execution by default, where this path doesn’t apply. But Accomplish’s deeper argument deserves attention: the kernel bug was the least interesting part. With AI-assisted vulnerability research industrialized — mainline-kernel privilege-escalation advisories running at roughly 2.3× the prior year, and public exploits landing within hours of fixes — a guest kernel is structurally always one bug behind. Containment that leans on the guest kernel being clean is containment that loses. Their four recommended locks (restrict unprivileged user namespaces, flip the seccomp filter to default-deny, disable module autoload for unused subsystems, and above all never share the whole host into the VM) each break entire categories of escape, not just individual CVEs.

Why this matters beyond the vendors

Three takeaways for anyone running coding agents in 2026:

The sandbox is the product. As Accomplish puts it, untrusted input isn’t an edge case for an agent — it’s the main case. The boundary between “the model did something silly” and “the model had my cloud credentials” is the only thing being sold here. When the harness itself runs commands outside the sandbox (git indexing, file watching), every one of those paths is attack surface the user never sees.

Patch velocity is a security feature — and it varies wildly. A week (Cursor, OpenAI) versus 50 days and ~30 releases (Anthropic) for the same class of bug is the kind of differential that should appear in enterprise procurement criteria. OpenAI says it is “continually strengthening our sandboxes, including tightening controls on where agents can write files.” Anthropic’s relative silence is itself data.

Model capability cuts both ways. Hiltch’s rhetorical question — “If these frontier models are so good, how come they’re not finding these critical vulnerabilities in their own products?” — is the uncomfortable one. The same models that industrialize vulnerability discovery for defenders industrialize it for attackers, and both sides now point at the same sandbox boundary.

For teams, the practical guidance is familiar but newly urgent: update Claude Code past 2.1.247, treat every cloned repo as hostile input, and if you’re running agents against anything sensitive, consider isolation that doesn’t depend on the agent’s own vendor sandbox — a VM whose host share is scoped, or an architecture where real credentials never enter the guest at all.

Accomplish, for its part, is betting the whole company on that last idea. This week, a lot of security teams are probably glad someone is.