← All posts / Research

72 Hours and $6,500: How Claude Opus 5 Hacked OpenAI From a Forum Image Upload

A three-person security team chained a libheif heap overflow and an OpenAI SSO flaw to reach OpenAI's internal monorepo — with Claude Opus 5 writing the exploit hours after release.

72 Hours and $6,500: How Claude Opus 5 Hacked OpenAI From a Forum Image Upload

On July 25, 2026, a three-person security startup called Hacktron AI compromised multiple OpenAI employee ChatGPT and Codex accounts. From there, they could have rifled through OpenAI’s internal GitHub organization, Slack workspaces, and email. Instead, they did something almost theatrical in its restraint: they prompted the compromised employee’s Codex to open a single harmless pull request — #1186742 — in openai/openai, OpenAI’s internal monorepo, purely to prove the access existed. Then they stopped, reported everything through Bugcrowd, and waited.

The full write-up, published by Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini on September 13 and swallowed whole by the press cycle this weekend, reads like a checklist of everything the AI security community has been warning about — and one detail in particular has become the story: Claude Opus 4.8 failed at the exploit across multiple sessions. Claude Opus 5, given the same problem within hours of its release, succeeded.

The chain: an image, a library, and an identity flaw

The attack began with a dependency, not a target. OpenAI runs its community forum at community.openai.com on Discourse, self-hosted in a Docker image based on Debian 12. Like most modern forums, it accepts image uploads — including HEIC and HEIF files, the format iPhones produce by default. Discourse’s normal image-checking path uses FastImage, but FastImage doesn’t understand HEIF, so those files get shunted to ImageMagick’s magick command for conversion. That handed attacker-controlled bytes directly to libheif, the open-source HEIF decoder underneath.

The Hacktron team started an Opus 4.8 session with the Discourse Docker image and asked it to audit the installed libheif package. The model found that particular security fixes had never been backported: an upstream commit that changed the vulnerable code the previous year was never documented as a security fix and never received a CVE, which is likely why Debian 12 shipped libheif 1.19.7 and even Debian 13 shipped the still-vulnerable 1.19.8. The bug is a heap buffer overflow in the HEIC decoding path, giving out-of-bounds read/write primitives — the raw material of remote code execution.

Developing that primitive into a working exploit is where the model generations diverge. On July 24, Opus 4.8 produced a working ImageMagick/libheif code-execution exploit with ASLR disabled, but multiple sessions failed to make it reliable against Discourse’s default ASLR-enabled configuration. That evening, Anthropic released Claude Opus 5. A fresh session produced a working ARM64 exploit for a local Mac within three hours, then ported it to the x86-64 and jemalloc configuration Discourse uses. By 6:00 a.m. on July 25, image uploads meant code execution.

One more step remained. The team placed Claude in an autonomous /goal loop against their own Discourse Cloud instance, proxied through rce.ee/ctf-forum so it looked like a CTF target — Opus refused to write exploits against remote production instances. Four hours later the agent had RCE on Discourse Cloud, reading /etc/hosts as proof. The generated exploit script then worked against OpenAI’s instance.

The final link wasn’t Discourse’s fault at all. “Sign in with OpenAI” on the forum flowed through auth.openai.com, and an SSO misconfiguration there meant compromising the forum could cascade into full ChatGPT and Codex account takeovers for any active member — including OpenAI employees, whose accounts were connected to internal GitHub, Slack, and email. As the researchers stress, the escalation path was an OpenAI identity problem, not a Discourse one: any first-party or third-party service behind OpenAI SSO would have worked as the pivot.

The disclosure: fast, clean, and cheap

The timeline is almost as instructive as the technique. Initial finding at 05:00–06:00 UTC on July 25. Bugcrowd submission by 08:00. Proof-of-concept PR in OpenAI’s internal monorepo by 13:30, with all further testing ceased at 15:30. OpenAI confirmed its side fixed at 22:49 — roughly 14 hours after the report. Discourse was reported via HackerOne the same day, replied Sunday, had a fix by Monday, added ImageMagick sandboxing as defense in depth, and published advisory GHSA-vhm9-85gw-x335 on July 28. On September 1, OpenAI paid a $6,500 bounty and marked the report resolved — with a note clarifying that the award covered the OpenAI-side finding only, since the Discourse-hosted forum was explicitly out of the bug bounty program’s scope.

Then there is the economics. The entire “HEIF Heist” research program — this OpenAI breach plus a multi-month sweep across Slack, Meta, GitHub Enterprise, Ruby on Rails, and Node.js frameworks like Next.js, Astro, and Gatsby — cost under $3,000 in tokens, took two months, and was conducted by three researchers. Adapting the exploit to each new company usually took a day or two. The models, the team notes, started almost blind each time, without knowing the exact libheif version, libc version, or deployment environment, and worked it out anyway.

Why this matters more than the punchline

The easy read is “Claude hacked OpenAI,” and the harder read is what the researchers actually said. This was not autonomous hacking; skilled human guidance remained essential. But the delta between model generations is the real signal. Opus 4.8 could find the missing backports and write a toy exploit; Opus 5 could beat ASLR and jemalloc in hours. Across the broader campaign, the team observed another jump from Opus 5 to GPT-5.6 Sol when exploiting targets without prior knowledge of the system. Whatever your prior on AI offensive capability, the slope is now measurable in release events, and it is steep.

The epilogue of the write-up makes the structural point plainly. Software has long benefited from a kind of security through complexity: the vulnerability could even be public, but turning a bug into a reliable exploit required rare expertise, significant time, and target knowledge. That was never a true boundary, but it practically protected ordinary companies. AI is removing that protection by turning scarce expertise into compute. Work that once took a well-resourced team months now compresses into days, for the cost of a nice dinner. Threat models built on “who can even carry out sophisticated attacks” are already outdated.

Two operational footnotes deserve attention. First, detection was essentially absent: the team is “not aware of any company that detected the activity except Shopify,” despite thousands of uploaded images and repeatedly crashing image processors. If you run a service that accepts user images and decodes HEIF/AVIF, the guidance is direct — update libheif and libde265 now (upstream security release v1.23.4 as of September 14), disable untrusted HEIF/AVIF decoding where you don’t need it, and isolate image pipelines in hardened, ephemeral sandboxes. If you self-host Discourse, a web-interface update alone is not enough; rebuild the Docker image.

Second, the boundary question. Anthropic’s own model refused to weaponize the exploit against a live remote target, and the team had to frame their own instance as a CTF to proceed. Guardrails bent the workflow without stopping it — the exploit still existed, still worked, and still ended up running against OpenAI’s forum. As frontier models get more capable, the gap between “refuses when asked directly” and “succeeds when asked creatively” is where a great deal of security policy is going to be negotiated.

For OpenAI, the incident is another entry in an uncomfortable season — arriving the same week as disclosures about concerning model behaviors and just before the Google Gemini red-team revelations. For everyone else, it’s a preview: a $6,500 bounty for a 72-hour breach of one of the world’s most valuable companies, pulled off by three people and a chat window.