Six to Zero: AISLE's Autonomous AI Finds 6 curl CVEs After OpenAI and Anthropic's Frontier Models Found None
Days after Anthropic Mythos and OpenAI Codex Security reported zero remaining flaws in curl, a startup's specialized AI system filed 29 reports — six became CVEs in curl 8.22.0, and the Linux kernel maintainer says he's seeing the same pattern.
On August 24, 2026, Daniel Stenberg — the founder and lead developer of curl — posted a status update that read like a clean bill of health. Only three CVEs were pending for the project’s next release, and the frontier AI cybersecurity systems he had been running against the codebase had come back empty. “[Anthropic] Mythos says it can’t find any more,” he wrote. “[OpenAI] Codex security shows an empty list.”
Within seventy-two hours, that narrative was upside down.
What happened
Stenberg has spent much of 2026 methodically throwing every serious AI code-auditing system at curl, one of the most heavily audited open-source codebases in existence. The library is deployed across an estimated 20 billion instances — from smart fridges to spacecraft — which makes even “Low severity” findings matters of real infrastructure consequence.
After the frontier models’ zero-result was published, a startup called AISLE ran its own autonomous vulnerability-discovery system against the same production code. The next day, before AISLE’s review process had even completed, Stenberg posted the first public comparison: “Mythos: 0, Aisle: 29.”
Of AISLE’s 29 reports, curl’s security team reviewed six within days and deemed them serious enough to merit public CVE designations in curl 8.22.0, which has now shipped:
- CVE-2026-80229 — OpenSSL provider use-after-free
- CVE-2026-80230 — OpenSSL pinning bypass
- CVE-2026-80231 — native CA store connection reuse
- CVE-2026-80255 — secure attribute bypass with tab
- CVE-2026-82208 — wolfSSL CA-cache hit overrides callback
- CVE-2026-82209 — domain-scoped public-suffix cookie
All six are rated Low severity. Three were reported on August 24, two on August 26, and one on August 27. By August 28, curl’s pending CVE count had risen from three to ten — six of the ten from AISLE alone.
Why the comparison is unusually clean
Head-to-head AI benchmarks are notoriously muddy. Models may have seen the answers in training data; evaluation harnesses reward benchmark-shaped behavior rather than real discovery. This episode has a property that almost no AI evaluation enjoys: the baseline was public and timestamped before the competing result existed.
Stenberg published the frontier systems’ zero-result on August 24. Only then did AISLE run its system against the live, production codebase. There was no capture-the-flag flag to capture, no benchmark with known answers — just current code, with curl’s maintainers (not AISLE) deciding what was real and what warranted a CVE. As AISLE’s founder Stanislav Fort put it, CVEs are imperfect markers, but each one represents a previously unknown flaw in production code, reproduced and accepted by domain experts, then fixed for deployed users.
Context: this fight has been running all year
The August result is the latest round in a running saga. Back in May, Stenberg documented an earlier audit where Anthropic’s restricted Claude Mythos model reported five “confirmed security vulnerabilities” — of which curl’s team validated only one low-severity issue after dismissing three false positives and one already-documented bug. At the time, experts were divided on whether Mythos’s showing meant frontier cyber capabilities were overstated or whether curl was simply running out of bugs to find.
The August zero-result leaned toward the second interpretation: perhaps the codebase was picked clean. AISLE’s 29 reports argue forcefully for the first — or at least for the claim that general frontier models and specialized discovery systems are now different classes of tool.
The Linux kernel echo
The pattern may not be confined to curl. Greg Kroah-Hartman, the longtime maintainer of Linux stable releases, responded to Stenberg’s post with a remark that deserves wider circulation: “I’m seeing the same for Linux as well. No idea what Aisle is doing differently, but wow…”
When the person who ships the world’s most deployed kernel says an AI system’s findings are leaving frontier-lab tools behind — and admits he can’t explain the gap — that is a signal worth taking seriously. It also raises a practical question for every large open-source project: if specialized AI systems can surface a 6x-to-29x report delta on the most audited code on Earth, what are they going to find in the median codebase, which enjoys none of curl’s scrutiny?
Why all six CVEs are “Low” — and why that’s the point
It’s tempting to dismiss Low-severity findings as noise. The opposite reading is more accurate. curl has absorbed decades of audits, fuzzing campaigns, and now a year of intensive AI scrutiny. The vulnerabilities that survive in such a codebase hide in narrow configurations and subtle interactions — an OpenSSL provider edge case, a cookie-domain scoping quirk, a TLS pinning bypass that only fires in specific builds. Finding them requires not brute-force pattern matching but patient, systematic enumeration of conditional paths.
That is precisely the kind of grinding, exhaustive work that specialized autonomous systems are built for, and that general-purpose models — optimized for breadth across a thousand tasks — may deprioritize. The severity ratings measure blast radius, not discovery difficulty.
The “System over Model” thesis
AISLE frames the result as evidence for its “System over Model” argument: that a purpose-built discovery system — tightly coupled tooling, sustained autonomous search, and domain-specific verification loops — can compete with and beat frontier-lab models at real-world zero-day discovery. One data point doesn’t prove a thesis, but the structure of this one is hard to argue with: same code, same week, public timestamped baseline, external adjudication by the maintainers themselves.
For the AI industry, the implication lands on both sides of the scale. Frontier labs should be concerned that their flagship security offerings are being publicly outperformed on their own benchmark territory. And security teams everywhere should note that the discovery gap between “best general model” and “best specialized system” is now measured in integer multiples, not percentages.
For curl and Linux users, the immediate news is simply good: six more real bugs found, confirmed, and fixed in curl 8.22.0. The deeper news is about who — or what — found them.
Sources
- AISLE: “AISLE Discovered Six curl CVEs After OpenAI and Anthropic Found Zero” (Sept 2, 2026)
- Daniel Stenberg: “Mythos finds a curl vulnerability” (daniel.haxx.se, May 11, 2026)
- Hacker News discussion: “Six curl CVEs after OpenAI and Anthropic came back with zero”
- SecurityWeek: “Claude Mythos Finds Only One Curl Vulnerability. Experts Divided on What It Really Means” (May 12, 2026)
- curl official security page: curl.se/docs/security.html
Sources
- [1] https://aisle.com/blog/aisle-discovered-six-curl-cves-after-openai-and-anthropic-found-zero
- [2] https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-vulnerability/
- [3] https://news.ycombinator.com/item?id=49536114
- [4] https://www.securityweek.com/claude-mythos-finds-only-one-curl-vulnerability-experts-divided-on-what-it-really-means/
- [5] https://curl.se/docs/security.html