← All posts / Models

OpenAI's GPT-5.6-Cyber Found Two Real Chrome Zero-Days

OpenAI's specialized cybersecurity model discovered two previously unknown V8 vulnerabilities, patched as CVE-2026-15903, and completes 95% of advanced security tasks.

OpenAI's GPT-5.6-Cyber Found Two Real Chrome Zero-Days

OpenAI has launched GPT-5.6-Cyber, a purpose-trained cybersecurity model that has already proven its mettle in the wild: it discovered two previously unknown zero-day vulnerabilities in Chrome’s V8 JavaScript engine, one of which was patched by Google as CVE-2026-15903. The model represents a significant escalation in AI-assisted vulnerability research — and a deliberate departure from the cautious refusal rates that have traditionally limited AI models in offensive security contexts.

The Discovery That Proved the Concept

The headline result is concrete and verifiable. OpenAI researchers used GPT-5.6-Cyber to investigate V8, the JavaScript engine that powers Google Chrome. The model surfaced two previously unknown vulnerabilities that could be chained together to corrupt memory and escape the V8 heap sandbox — a class of bug that sits at the intersection of compiler theory and exploitation craft.

The more severe of the two, designated CVE-2026-15903, is a high-severity flaw in V8’s optimizing compiler. The compiler incorrectly skipped a safety check, creating a condition that an attacker could leverage to achieve out-of-bounds memory access. When chained with the second vulnerability, an attacker could potentially break out of V8’s heap sandbox — a critical step toward full browser compromise.

Google received the disclosure through coordinated vulnerability disclosure and shipped a fix. The fact that a single AI model, operating in a research capacity, independently surfaced novel bugs in one of the world’s most heavily fuzzed codebases is a watershed moment for AI-driven security research.

95% Completion Rate: A Deliberate Alignment Shift

Perhaps the most striking number from the launch is the completion rate. OpenAI built an internal benchmark it calls the Advanced Cybersecurity Tasks evaluation. On this test, GPT-5.6-Cyber completes 95.0% of requests. That figure stands in stark contrast to its predecessors:

  • GPT-5.5-Cyber (the immediate predecessor): 57.3% completion
  • GPT-5.6 Sol (the standard, non-cyber variant): 1.5% completion
  • GPT-5.6 with Daybreak Blue access: 2.0% completion

The jump from 57.3% to 95% represents a generational improvement, but the gap between the cyber-specialized model (95%) and the standard model (1.5%) is even more telling. Standard GPT-5.6 refuses nearly all dual-use security requests — building exploits, analyzing malware, constructing attack chains. GPT-5.6-Cyber was specifically trained to reduce those refusals for authorized users while maintaining guardrails against misuse.

OpenAI also reported that the model discovered 3× more exploits than its predecessor and significantly cut false positives, making it not just more willing but more accurate.

Daybreak: A Two-Tier Access Architecture

The model is not generally available. It lives inside OpenAI’s Daybreak program, which was expanded alongside this launch with a two-tier access structure:

Daybreak Red provides access to GPT-5.6-Cyber and other purpose-trained models for authorized vulnerability research, exploit validation, and offensive security work. This tier is restricted to vetted defenders and security researchers who have gone through an approval process.

Daybreak Blue provides a more constrained set of capabilities oriented toward defensive tasks — secure code review, vulnerability identification without exploit development, and malware analysis at a higher level. Even within Blue, the standard GPT-5.6 model completes only about 2% of advanced cybersecurity task requests, reflecting the same alignment-driven refusals that limit general-purpose models.

The tiered approach is OpenAI’s answer to a fundamental tension: the same capabilities that make a model effective at finding vulnerabilities also make it dangerous in the wrong hands. Rather than hobbling the model universally, OpenAI has opted for capability stratification gated by access controls.

The Cyber Defense Window Is Closing

The launch post is titled “Expanding Daybreak as the Cyber Defense Window Narrows,” and that framing is deliberate. OpenAI’s argument is that AI-powered offensive tools will inevitably become accessible to adversaries. The question is not whether attackers will use AI to find zero-days, but whether defenders will have equivalent or better tools first.

GPT-5.6-Cyber is positioned as the answer — a model that can find vulnerabilities and construct proof-of-concept exploits before attackers do. The V8 discovery is the proof point: a bug that existed in Chrome’s most critical engine, found not by a human researcher but by an AI model operating semi-autonomously.

This framing has implications for the broader security ecosystem. If AI models can reliably surface novel vulnerabilities in heavily audited code, the economics of bug hunting shift dramatically. The cost of finding a high-severity zero-day — long the domain of elite researchers — may be falling rapidly.

What GPT-5.6-Cyber Does Not Do

It is worth noting the caveats. The 95% figure measures completion rate, not accuracy. A model that completes 95% of requests is not necessarily producing correct or useful output 95% of the time — it simply refuses less often. Independent reporting from The New Stack noted that GPT-5.6-Cyber “did not outperform Sol across the board” and that its reports “tended to be shorter and less detailed” in some categories.

The model is also narrowly scoped. It excels at vulnerability research and exploit validation but is not a general-purpose security platform. And despite the reduced refusal rate, OpenAI maintains that the model is better at finding and fixing vulnerabilities than at weaponizing them — an assertion that remains difficult to verify independently given the access restrictions.

Industry Context and Implications

The launch positions OpenAI squarely in the cybersecurity AI race, alongside companies like Google (Project Zero’s AI experiments), Microsoft (Security Copilot), and a growing field of startups building AI-driven vulnerability scanners. But GPT-5.6-Cyber is distinct in its willingness to handle offensive capabilities — exploit development and chain construction — that most competitors shy away from.

The Chrome zero-day discovery gives OpenAI a credibility marker that marketing alone cannot buy. Finding a novel bug in V8 is hard. Finding two that chain together is harder still. The fact that the research resulted in a real CVE and a real Google patch means the model’s output survived external validation — the bugs were real, the fixes were real, and the coordinated disclosure worked.

For the security community, the message is clear: the tools are arriving. Whether the defense window is truly narrowing or simply being redefined is a question that the next 12 months of AI-driven vulnerability research will answer.