The Frontiers Refuse to Fight: Inside Artificial Analysis's New Cyber Defense Index
The new Artificial Analysis Cyber Index benchmarks AI on the full defensive loop — find, reproduce, patch — across 351 expert-vetted tasks. The twist: frontier models refuse up to 98% of memory-safety tasks, leaving Grok 4.7 and Xiaomi's MiMo-V2.6-Pro tied at the top.
Benchmarking shop Artificial Analysis has quietly launched the most serious attempt yet to answer a question every security team is asking: which AI model should I trust to find and fix vulnerabilities in my code? The answer, revealed by the new Cyber Index, is not the one anyone expected.
Announced September 25 alongside the newly formed Cyber Index Alliance — with Collinear AI, IBM, NVIDIA, and Vercel as launch partners — the index benchmarks models on the complete defensive loop: discovering vulnerabilities in a codebase, reproducing and validating them, and patching them without breaking existing functionality. Crucially, it is defense-only. Models work from source code the way a security engineer auditing an application would, and no model is ever asked to build a working exploit.
The headline result is a paradox. On the overall index, Grok 4.7 (xhigh) and Xiaomi’s MiMo-V2.6-Pro are tied at the top with scores of 56, followed by GPT-6 Luna (max) at 53. Not Anthropic’s Claude. Not the most expensive frontier tiers. And the reason is structural: on the memory-safety evaluation, GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B, and Qwen3.8 27B all refuse at least 98% of tasks, making it impossible to even assess frontier model performance on an entire class of real-world vulnerabilities.
Three benchmarks, one defensive loop
The Cyber Index v1 is an equally weighted average of three evaluations, totaling 351 tasks vetted by industry partners and academia:
CWE-Bench-AA (from Collinear AI) is a full security audit. The agent receives a checkout of an open-source repository and an instruction to audit the code and fix what it finds. Tasks reproduce disclosed CVEs, describe the area of concern but not the exact location, and span all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. The test set contains 120 private, held-out tasks. A task counts as solved only when a programmatic verifier confirms the exploit no longer works and legitimate behavior still works — no partial credit, no LLM judge.
DeepsecBench-AA (from Vercel) isolates the discovery process. The agent reviews scanner-flagged files in open-source application code and reports every vulnerability it can confirm, scored via F2 (recall weighted above precision) against a golden set of expert-verified findings.
CyberGym-E2E-AA (from Berkeley RDI) is the hardest of the three: 131 tasks, one per project, targeting real memory-safety vulnerabilities in widely used C/C++ projects such as FFmpeg and CPython. The model must locate the bug, write a proof-of-concept input that triggers the crash, and patch it so the crash no longer reproduces — while the project’s own functionality tests still pass.
All three run on Stirrup, Artificial Analysis’s open-source agent harness, in sandboxed environments with no internet access. And because cyber work is dual-use, the team tracks safety refusals separately from scores — a declined task scores zero.
Why the frontier refuses
The refusal finding deserves its own headline. Memory-safety work — the C/C++ vulnerability class behind a large share of critical CVEs — sits squarely in the dual-use zone where providers’ safety systems intervene. Asking a model to write a proof-of-concept input that crashes FFmpeg looks, to a filter, indistinguishable from asking it to craft an exploit.
The consequence is a benchmark that cannot measure the best models on an entire category of defensive work. Frontier labs argue their refusals are the responsible default; security practitioners counter that an agent that won’t look for a null-pointer crash is useless to the teams defending infrastructure from it. The Cyber Index makes this tension visible and quantified for the first time — the refusals are right there in the safety-blocks column, model by model, benchmark by benchmark.
How models actually fail
The failure analysis is where the index earns its keep. On CWE-Bench-AA, failures concentrate in two patterns: partial fixes (55% of failed attempts — the model fixed the primary issue but left a related one open, such as a second entry point) and over-corrections (~24% — the patch breaks legitimate behavior). Tellingly, over-correction is most common among the strongest models, accounting for roughly 40% of failures for the four highest scorers versus ~15% for the lowest, because confident models keep editing until they break an edge case of legitimate functionality.
On DeepsecBench-AA, the best model identifies only 41% of expert-verified issues. Models reliably find flaws with a direct path from untrusted input to consequence, but rarely find bugs requiring reasoning through a sequence of events or business and privacy rules. When models do report those harder sequence-of-events vulnerabilities, they are almost always right — 95% of such reports are correct — and GPT-6 Sol and GPT-6 Astra find them in roughly 30% of runs, two to three times the rate of typical models. That, Artificial Analysis suggests, may be an emerging capability.
On CyberGym-E2E-AA, discovery is the bottleneck: 42% of attempts hit the 90-minute limit without ever producing a crashing input. Models pass 50% of attempts on out-of-bounds bugs, but only 33% on use-after-free (which depends on reasoning about an object’s lifetime across operations) and 20% on integer/arithmetic bugs. And in a finding that should humble any agent maximalist: 31% of passing attempts fixed the wrong bug — a real crash, but not the target vulnerability, usually a shallower issue like a null-pointer dereference the model stopped at because it was the first thing it could validate.
The Alliance and what comes next
The Cyber Index Alliance is structured as a standing consortium: partners contribute expert input on scope and methodology, and may contribute datasets directly. Collinear AI contributed CWE-bench as a private held-out evaluation; Vercel did the same with DeepsecBench; IBM and NVIDIA advise on methodology. Organizations can join via cyber@artificialanalysis.ai.
The roadmap is explicit about gaps: incident response, writing new code without introducing vulnerabilities, and targets without source access (compiled software, live servers) are all planned future additions. Exploit realization — turning a found vulnerability into a working exploit — is permanently out of scope for a defense-focused index.
Why it matters
Until now, “cyber capability” in model marketing has meant offensive-flavored demos or single-metric scores on synthetic CTF puzzles. The Cyber Index instead measures the boring, decisive work: can an agent audit a real repository, find the bug a scanner flagged, prove it crashes, and patch it without collateral damage — for a cost a security team would actually pay? The cost-per-task axis, plotted against score, is arguably the most enterprise-relevant chart Artificial Analysis has published.
The first leaderboard already delivers a verdict on the industry’s safety posture: the models most capable of defending memory-unsafe code are, by their makers’ own configuration, the ones least willing to try. Whether that balance is right is now a question with numbers attached — and a consortium with every incentive to keep asking it.
Sources
- [1] https://artificialanalysis.ai/articles/artificial-analysis-cyber-index
- [2] https://artificialanalysis.ai/evaluations/artificial-analysis-cyber-index
- [3] https://www.developer-tech.com/news/artificial-analysis-cyber-index-tests-defensive-ai-security-models/
- [4] https://news.lavx.hu/article/artificial-analysis-launches-cyber-index-for-ai-powered-cyber-defense