← All posts / Tools

Anthropic Flips the Default: Claude Code's Auto Mode Replaces Human Permission Prompts

Auto mode is now the default in Claude Code for Pro, Max, and Team plans — classifier-gated autonomy that blocked 89% of dangerous commands in testing while fatigued humans caught just 13.6%.

Anthropic Flips the Default: Claude Code's Auto Mode Replaces Human Permission Prompts

For years, the safety story of AI coding agents has rested on a ritual: before the agent runs a command or touches a file, a human clicks “approve.” On August 7, 2026, Anthropic published the results of months of testing that undermine that ritual — and then acted on them. Starting August 14, new sessions in Claude Code on Pro, Max, and Team plans run in auto mode by default, replacing routine permission prompts with a model-based classifier that decides, per tool call, whether an action is safe to execute. The Enterprise, API, and cloud-platform rollout is planned within the month, with admins notified before the switch.

The change is more consequential than a settings toggle. It is the first time a major AI lab has looked at its own permission system, measured how humans actually behave inside it, and concluded that the human in the loop was the weakest part of the loop.

The numbers that justified the flip

Anthropic’s case rests on a blunt statistic: Claude Code users approve 97% of permission prompts. Most prompts are for routine, safe commands — but a 97% approval rate means many users are clicking through reflexively rather than reviewing. When Claude presents a plan for approval, users reject 39% of them; for individual permission requests, rejection is just 3%. The pattern shows where attention actually goes, and it isn’t the permission dialog.

Users had already been voting with their config files. As of June 2026, 49.5% of active CLI users had manually created a Bash allow-rule — 5% allow any shell command outright, and another 43% maintain interpreter rules like Bash(python:*) that are functionally equivalent. That share grows roughly 5 percentage points every five weeks. Meanwhile, 62% of users have used bypassPermissions or clicked “don’t ask again” on Bash, and 25% of interactive sessions start in bypass mode outright. The permission system wasn’t holding the line; users were routing around it in droves.

The decisive experiment: Anthropic hired 1,053 paid professional testers through a research vendor for a coding study. Partway through each session, a single permission prompt was quietly swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it. The humans caught the dangerous command 13.6% of the time (143 of 1,053). Auto mode blocked 89% of the same commands (937 of 1,053). Head to head, the classifier blocked 800 commands that a human approved; humans blocked only 6 that auto mode allowed. And humans degraded with fatigue — blocking about 17% of dangerous commands early in a session, dropping to about 5% after 50 or more prior prompts — while the classifier’s block rate stayed flat regardless of session length.

Real-world session data pointed the same direction. Examining safety-flagged sessions from May–June 2026 on two 0–10 harm-severity scales, Anthropic found that manually approved sessions contained serious unintended harm more than twice as often as auto mode sessions: at production-level severity (7+), 6.3% of manually approved sessions included a harmful action the user hadn’t explicitly requested, versus 2.4% for auto mode.

What auto mode actually is

Auto mode is not “skip permissions with extra steps.” It is a two-layer pipeline, documented in an Anthropic engineering post from March 2026 that explains the design in unusual detail.

At the input layer, a server-side prompt-injection probe scans tool outputs — file reads, web fetches, shell output, external tool responses — before they enter the agent’s context. When content looks like an attempt to hijack the agent’s behavior, a warning is injected telling the agent to treat the content as suspect and anchor on what the user actually asked for.

At the output layer, a transcript classifier (running on Sonnet 4.6) evaluates each pending action before it executes. It runs in two stages: a fast single-token yes/no filter, followed by chain-of-thought reasoning only when the first stage flags something — so reasoning tokens are spent only where needed. By design, the classifier sees only user messages and the agent’s tool calls, with Claude’s own reasoning stripped out, making it “reasoning-blind” — it judges actions, not the agent’s justifications for them.

The threat model is explicit. Agents take dangerous actions for four reasons: overeager helpfulness, honest mistakes about blast radius, prompt injection, and (still theoretical in practice) genuine misalignment. The classifier is tuned to the first two, but the defense — block the action — applies to all four. The blocked-action taxonomy reads like an incident log because it is one: scope escalation (“clean up old branches” leading to mass remote-branch deletion), credential exploration (grepping env vars for alternative tokens after an auth error), agent-inferred parameters (deleting the wrong “closest match” job), data exfiltration via public Gist, and safety-check bypass (retrying a failed deploy with a skip-verification flag).

Crucially, auto mode has a failure posture. When the classifier blocks something, Claude usually finds a safer path on its own or asks the user directly; if it can’t make progress — three blocks in a row, or twenty across a session — Claude Code falls back to manual approvals. And broad allow rules that would grant arbitrary code execution (like python:*) are set aside while in auto mode, since they would let commands skip the classifier entirely.

Hardened against adversaries, tested against competitors

Anthropic also worked with Apollo Research on a two-week adversarial pilot, injecting synthetic attacks into real coding trajectories. After hardening, re-testing on a held-out attack set Anthropic never saw caught about half of previously missed attacks at the cost of about 2% of previously caught ones — evidence the fix generalized rather than overfit. Across shared and held-out sets, the classifier’s miss rate fell from 12% to 7% (measured against deliberately malicious synthetic attacks, not real traffic).

The most eyebrow-raising external data came from Trajectory Labs, a third-party evaluator that tested 72 indirect prompt-injection scenarios, each run 10 times, using attacks optimized against a model that wasn’t in the reported set. Zero of 720 attack attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode. By comparison, GPT-5.6 Sol in Codex’s Auto-review mode fell to 5.83% of attacks, and 19.03% in Codex Full Access mode. Anthropic notes the evaluation measured the underlying models via an identical third-party browser integration, not the full first-party safeguards of either product — but as a controlled comparison, it is striking.

What it means in production

The productivity argument is quantified too: among Teams and Enterprise adopters, auto mode users ship about 25% more PRs, because agents run longer without interruption. Anthropic itself runs auto mode as the default for all internal Claude Code usage, and published three incidents the classifier caught internally: a failed internal upload whose fallback would have published a report to a public code-sharing site; a mass process kill across roughly 2,000 pods that would have torn down hundreds of GPUs of in-flight training jobs; and a read-only billing-access request where Claude reached for an account-root admin role.

Named customers are already living with the new default. Adobe’s merchandising platform team runs an agentic loop in auto mode to build and verify pricing pages across 90+ countries and 30+ languages. Nuro uses it for overnight research agents that hill-climb evaluation metrics and return finished PRs by morning — one engineer described kicking off an agent at 10 p.m. and receiving three PRs at 5 a.m. Gusto adopted it explicitly to end the permission fatigue that was pushing engineers toward bypassing checks entirely; about 10% of its sessions since mid-May include a classifier denial. Garner Health pushed auto mode to all 550 employees via managed settings, standardizing an SDLC that no longer depends on hand-curated command allowlists.

The honest caveats

Anthropic’s own framing is careful: auto mode “reduces risk for most users” but “relies on classification systems and therefore does not eliminate risk,” and the company still recommends manual review for high-stakes production changes. Critics have made a subtler point: auto mode isn’t competing against a theoretical ideal of deterministic sandboxing — it’s competing against what developers actually do, which is approve nearly everything and gradually disable the prompts. Judged against that baseline, the classifier wins on the data. Judged against the ideal of provable guarantees, it’s a probabilistic guardrail replacing a worn-out human one.

There’s also a business logic worth naming. Classifier overhead — a small number of extra tokens per tool call — is now free for Pro, Max, and Team users, and Anthropic plans to stop charging for it everywhere. Every hour an agent runs autonomously is an hour of tokens billed; unblocking long-running work with Opus 5-class models isn’t just a safety decision, it’s a revenue decision. Both things can be true at once.

The default flip is the real news, though. The industry spent 2025–2026 debating whether humans should supervise AI agents more or less. Anthropic has now run the largest controlled study of the question inside a production coding tool, didn’t like the answer it got about human supervision, and changed the default accordingly. Every agentic IDE, terminal agent, and autonomous-coding startup now has to answer the same question: if your permission prompts are approved 97% of the time, are they protecting anyone — or just performing oversight?

This article is based on Anthropic’s official announcements and engineering documentation, TechCrunch and The Decoder reporting, and third-party evaluation data cited by Anthropic.