← All posts / Meta

Rival's Weapon, Your Crown Jewels: Hacktron Used Claude Opus 5 to Hack OpenAI in 72 Hours

A three-person team chained a libheif heap overflow in OpenAI's community forum into ChatGPT/Codex account takeovers and an internal monorepo PR — with Claude Opus 5 writing the exploit. The age of cheap, AI-driven exploitation has arrived.

Rival's Weapon, Your Crown Jewels: Hacktron Used Claude Opus 5 to Hack OpenAI in 72 Hours

On July 25, 2026, a three-person security research team called Hacktron AI obtained remote code execution on OpenAI’s own community forum, pivoted through a flaw in OpenAI’s single sign-on infrastructure, and took over the ChatGPT and Codex accounts of OpenAI employees. To prove the breach without touching anything sensitive, they prompted a compromised employee’s Codex to open a pull request — PR #1186742 — in OpenAI’s internal monorepo, openai/openai.

The full chain, from first probe to internal repository access, took less than 72 hours. The exploit development that made it possible was largely written by Anthropic’s Claude Opus 5 — the flagship model of OpenAI’s most prominent rival. The story went wide this week when the Wall Street Journal, The Information, the Guardian, Forbes, CBS, and the Financial Times all picked it up within a single news cycle, and it stands as the sharpest illustration yet of a uncomfortable new reality: AI has collapsed the cost of sophisticated exploitation.

The kill chain, step by step

The attack is a textbook dependency-and-identity chain, and every link is instructive.

1. A forum with a soft underbelly. OpenAI runs its developer community at community.openai.com on Discourse, the open-source forum platform, and lets users “Sign in with OpenAI” through auth.openai.com. Hacktron’s insight — that compromising the forum could create a path into broader OpenAI services through this identity flow — was the strategic key. But Discourse itself is a hardened target the team had studied before. So they went after a dependency instead.

2. An image decoder nobody watches. Discourse’s image-upload pipeline uses FastImage for most formats, but HEIC and HEIF files — Apple’s default photo format — follow an unusual path: FastImage doesn’t understand them, so they get handed to ImageMagick’s magick command for conversion. That exposes the underlying libheif parser directly to attacker-controlled files.

3. A patch that never arrived. On July 23, the team pointed an Opus 4.8 session at the Discourse Docker image and asked it to audit the installed libheif package. The model found that specific upstream security fixes had not been backported — the vulnerable code had actually been changed upstream the previous year, but the commit was never documented as a security fix and never received a CVE. As a result, Debian 12 and 13 both shipped vulnerable versions (1.19.7 and 1.19.8 respectively), and Discourse’s Docker image, built on Debian 12, inherited the flaw. The bug allows a heap buffer overflow during HEIC decoding, yielding out-of-bounds read/write primitives.

4. Opus 4.8 stalls; Opus 5 delivers. On July 24, the team used Opus 4.8 to develop a working ImageMagick/libheif code-execution exploit — but only with ASLR disabled. Multiple sessions trying to make it reliable against Discourse’s default ASLR-enabled configuration went nowhere. That evening, Anthropic released Claude Opus 5. A fresh session produced a working ARM64 exploit for a local Mac within three hours, then ported it to the x86-64 and jemalloc configuration Discourse uses. By 6:00 a.m. on July 25, they had confirmed local RCE through an image upload. They then placed Claude in an autonomous /goal loop against their own Discourse Cloud instance — proxied through rce.ee/ctf-forum so it looked like a CTF target, because Opus refused to write exploits against real remote systems. By 10:00 a.m., the agent had achieved remote RCE and demonstrated it by reading /etc/hosts. With the generated exploit script in hand, they popped OpenAI’s instance.

5. The SSO flaw that turned a forum into a skeleton key. Here is the part that should keep every platform team awake: the escalation was not Discourse-specific. An OpenAI SSO misconfiguration meant that compromising any service in the identity perimeter — the forum being merely one proof — yielded takeover of the ChatGPT and Codex accounts of anyone who logged in, with no interaction from the victim. Because people connect GitHub, Slack, and email to Codex and ChatGPT, the theoretical blast radius included OpenAI’s source-control systems. The team took over one employee’s account, whose Codex was connected to OpenAI’s GitHub organization, and used it to file the harmless PR. Then they stopped, and reported.

The economics are the story

Hacktron’s disclosure timeline shows responsible handling throughout: a Bugcrowd submission within hours of confirming cross-product impact, proof-of-concept and direct alerts the same day, OpenAI confirming a fix roughly 14 hours after the initial report, Discourse shipping a fix over the weekend and adding ImageMagick sandboxing as defense in depth, and a public advisory (GHSA-vhm9-85gw-x335) on July 28. OpenAI paid a $6,500 bounty on September 1 — with a note that the forum itself was technically out of scope, and the award recognized the OpenAI-side finding.

But the number that should reframe your threat model is this: the broader “HEIF Heist” campaign — auditing libheif across Slack, Meta, GitHub Enterprise, Ruby on Rails, and Node.js frameworks like Next.js, Astro, and Gatsby — took two months, three researchers, and less than $3,000 in tokens. Adapting the exploit to each new company typically took the AI one or two days, working nearly blind, without knowing the exact libheif version, libc, or deployment environment. Only Shopify noticed anything, even after thousands of uploaded images repeatedly crashed image processors.

The team also reports a clear capability gradient across model generations: Opus 4.8 struggled across several sessions to build a reliable ASLR-enabled exploit; Opus 5 succeeded within hours; and GPT-5.6 Sol later showed another jump when exploiting targets about which almost nothing was known. Skilled human guidance remained essential — this was not fully autonomous hacking — but the amount of work a three-person team could do expanded dramatically.

Security through complexity, gone

The deepest lesson in Hacktron’s epilogue is about a boundary that quietly protected almost everyone. Software has long relied on “security through complexity”: the code and even the vulnerability could be public, but weaponizing a memory-corruption bug into a reliable exploit demanded rare expertise, target-specific knowledge, and months of effort. Zero-days were reserved for the highest-value targets precisely because they were expensive. AI is dissolving that economic shield by “turning more of this scarce expertise into compute.” Work that once required a well-resourced team and months now compresses into days, at a cost measured in hundreds of dollars of API credits.

The defensive playbook follows directly. Update libheif and libde265 through your distribution’s security channel — as of September 14, 2026, the latest upstream security release is v1.23.4. Better yet, disable untrusted HEIF/AVIF decoding where you don’t need it, and run image-processing pipelines in hardened, ephemeral sandboxes; ImageMagick’s security policy can restrict accepted formats and resource usage. And treat your SSO perimeter as a single, critical asset: any service inside it is now a bridge to everything else, and the forum you forgot about may be the cheapest way in.

There is a bitter symmetry here worth sitting with: the model that broke OpenAI’s forum belonged to Anthropic, and the next capability jump the researchers measured came from OpenAI’s own GPT-5.6. Frontier labs are sharpening both sides of the blade faster than most organizations are updating their threat models. The question Hacktron’s PR leaves behind isn’t whether AI-assisted exploitation works — it’s who else has been quietly running the same playbook, without a bounty payout and a blog post at the end.