← All posts / Research

One Model Finished the Hack: Booz Allen's Cyber Weapon Index Ranks 18 AIs as Attackers — and a Cheap Harness Erases the Ranking

Booz Allen put 18 US and Chinese AI models against a live corporate network as autonomous attackers. Only Claude Mythos completed the full kill chain — then 15th-place Claude Sonnet 5 matched it once given an attack harness.

One Model Finished the Hack: Booz Allen's Cyber Weapon Index Ranks 18 AIs as Attackers — and a Cheap Harness Erases the Ranking

Booz Allen Hamilton has published its first-ever Cyber Weapon Index (CWI), and it is the most sobering AI security benchmark released to date. The consulting and defense technology firm put 18 leading large language models — nine American, nine Chinese — into the role of autonomous attackers, each controlling a real attacker machine against a production-grade enterprise network, and simply told them to break in. No tool menu. No supporting scaffolding. No human steering.

Only one model finished the job: Anthropic’s Claude Mythos. And in a twist that undercuts the entire league table, a 15th-place model nearly matched it once the testers bolted on a cheap piece of supporting software. The report’s conclusion is blunt: the model is no longer the unit of risk. The system is.

How the test worked

Most AI cyber benchmarks measure what a model knows — quiz-style evaluations of vulnerability trivia and capture-the-flag puzzles. The CWI measures what a model does. Each of the 18 models was given its own attacker machine and issued commands one at a time against a defended Active Directory network, with the goal of progressing from initial access all the way to full domain administrator control.

Three design choices make the results unusually credible:

  • No curated tools. Models received no menu of hacking utilities or agentic scaffolding, isolating what the raw model can accomplish alone.
  • Logs over claims. Every action was validated through network telemetry, host logs, domain controller data, and intrusion-detection sensors. Points were awarded for what the network logs proved, not what the model claimed it did.
  • Symmetric conditions. US and Chinese models ran under identical conditions, enabling the first apples-to-apples offensive-capability comparison across the two AI ecosystems.

Each model’s CWI score combines a vulnerability research score (VRS) — can it identify planted or novel vulnerabilities in compiled software without source code — and a kill chain attainment score (KCAS), which awards points based on how far the model progresses through an end-to-end intrusion, tested both with and without stolen credentials.

The scoreboard

The 18 models ranked from highest to lowest CWI score:

  1. Claude Mythos (Anthropic) — 80
  2. Grok-4.5 (xAI) — 49
  3. GPT-5.6 Sol (OpenAI) — 46
  4. Muse Spark 1.1 (Meta) — 38
  5. Kimi K3 (Moonshot AI) — 38
  6. GLM-5.2 (Z.ai) — 37
  7. Claude Opus 4.8 (Anthropic) — 36
  8. GPT-5.5-Cyber (OpenAI) — 34
  9. Nemotron-Ultra (Nvidia) — 33
  10. DeepSeek-V4-Pro — 23
  11. DeepSeek-V4-Flash — 17
  12. Qwen3.5-397B (Alibaba) — 17
  13. MiniMax-M3 — 15
  14. Nemotron-Super (Nvidia) — 15
  15. Claude Sonnet 5 (Anthropic) — 13
  16. GLM-4.5-Air (Z.ai) — 11
  17. Qwen3.6-35B (Alibaba) — 9
  18. Qwen3-Coder (Alibaba) — 4

Claude Mythos’s performance stands apart. Handed a stolen employee credential, it took administrator control on every attempt — and it independently worked out its own route to higher privileges based on what it found inside the network, rather than following any predetermined attack plan. On the harder test, with no credentials at all, it broke in from the outside and still achieved full domain compromise. No other model managed that.

The capability distribution beneath the leader is what should worry CISOs:

  • Four models — Mythos, Grok-4.5, Muse Spark 1.1, and GLM-5.2 — reached full domain access and control.
  • Four more — GPT-5.6 Sol, Kimi K3, GPT-5.5-Cyber, and DeepSeek-V4-Pro — achieved lateral movement across the network.
  • Two — Claude Opus 4.8 and Qwen3.5-397B — obtained credentials that let them expand access and privileges.
  • All but one (Qwen3-Coder) autonomously gained initial access to the network.

Notably, Booz Allen found no substantial differentiation between US and Chinese models. The capability curve is a function of frontier training scale, not geography — and the firm’s assessment is that most of the 18 will arrive at Mythos’s full kill chain capability within six months.

The finding that erases the ranking

Then the report undermines its own league table. Claude Sonnet 5 placed 15th of 18 with a score of 13. When Booz Allen paired it with an attack harness — the software layer that connects a model to hacking tools and keeps it on task — it rivalled Mythos. A 67-point gap, closed by plumbing.

A harness lets a model stay focused, adapt, recover from failures, and chain single actions into a sustained operation. The implication is profound for threat modeling: an adversary doesn’t need a frontier model. They need a mid-tier model and good engineering. Booz Allen concedes it has not tested Chinese or open-weight models paired with optimized harnesses — and its results “strongly suggest” such fully capable combinations already exist in the wild.

A second underreported finding reinforces the point. One model refused a task on the grounds that it had no credentials. Its cyber-tuned sibling, handed the identical task, complied and carried it out. Booz Allen’s general rule: guardrails are not a fixed property of a model. Their effectiveness shifts with context and configuration. A refusal observed in one setting tells you nothing about another.

Where every model still fails

The clearest limit is real-world vulnerability research. Against a deliberately planted flaw, every model — American, Chinese, open-weight, closed — scored near the ceiling. Provenance made no difference.

Against a genuine, unseen vulnerability buried in a large production library, all nine frontier API models scored zero. One unnamed leading model correctly analyzed the vulnerable component, then dismissed it as safe. Only Anthropic’s frontier models spotted the flaw at all, and only Mythos understood it well enough to exploit it. (Mythos has form here: it reportedly found 10,000 critical vulnerabilities in a single month in May.)

Booz Allen frames this gap as breathing room: real-world offensive capability still trails benchmark performance, buying defenders time before the gap closes. The firm calls mainstream AI-enabled attacks — from ransomware gangs and state-backed groups alike — “imminent,” citing July’s Hugging Face breach as the moment a model first completed the kill chain in the real world rather than a laboratory.

What Booz Allen wants Washington to do

The report’s policy asks read like a defense contractor’s wishlist — and should be read alongside the fact that Booz Allen launched Vellox Labs Guile, a counter-AI product, in the same release (claiming, unverified by any third party, that coordinated counter-AI playbooks cut attacker success by over 95% in its own testing). But the asks themselves are separable from the sales pitch:

  • Binding readiness standards: enforceable, sector-specific deadlines forcing critical infrastructure operators to prove they can contain an AI-enabled intrusion.
  • Continuous measurement: a national program testing foreign and open-weight models under realistic conditions, since adversaries will pair them with optimized harnesses.
  • Offensive overmatch: “We must aggressively develop agentic capabilities that accelerate authorized offensive cyber operations while simultaneously building AI-enabled defenses that detect, decide, and respond at machine speed,” the report argues.
  • Governed access for vetted defenders to the very capabilities they are meant to defend against.

One model is conspicuously absent from the index: OpenAI’s GPT-6 Astra, which was not tested. OpenAI said on Tuesday that Astra has already reached the company’s “critical” cybersecurity capability threshold — meaning it is so capable at finding and exploiting zero-day bugs that it poses significant risk both from malicious users and from the model itself, should it be misaligned. The index, in other words, arrives measuring a field that has already moved.

The bottom line from McLean: the model is now the attacker. The question is no longer whether AI can execute sophisticated intrusions autonomously — one demonstrably can, seventeen more are close behind, and a well-built harness can make even the stragglers dangerous.