← All posts / Tools

ChatGPT Work Meets the Lethal Trifecta: Simon Willison's Four-Hours-Long Security Autopsy of OpenAI's Agent Platform

Security researcher Simon Willison spent weeks reverse-engineering ChatGPT Work and published his findings August 30: an open-internet code sandbox, a full headless Chrome, a persistent shared filesystem, Cloudflare Workers deployment, sub-agents, and scheduled prompts — every ingredient of his 'lethal trifecta' attack model, combined in one product.

ChatGPT Work Meets the Lethal Trifecta: Simon Willison's Four-Hours-Long Security Autopsy of OpenAI's Agent Platform

On July 9, 2026, OpenAI announced ChatGPT Work — and then, in security researcher Simon Willison’s words, began “furiously iterating” on it. Nearly two months later, the product had accumulated so many capabilities so quickly that nobody outside OpenAI seemed to have a clear picture of what it actually was. On August 30, Willison published “Understanding ChatGPT Work,” a nearly 2,000-word teardown based on weeks of hands-on reverse-engineering. His conclusions matter for anyone whose company is standardizing on ChatGPT: the platform now bundles every ingredient of the attack model he calls the “lethal trifecta” — private data, untrusted content, and exfiltration paths — in a single consumer-reachable product, and OpenAI’s documentation still does not explain how it defends against prompt injection.

Two products wearing one name

Willison’s first finding is taxonomic. “ChatGPT Work” is actually two distinct products. The first — which he dubs Work Cloud — runs on OpenAI’s infrastructure, accessible from chatgpt.com and the mobile apps. The second — Work Local — lives inside the ChatGPT desktop app, the application formerly known as Codex, and can access files and run programs directly on your computer. He characterizes Work Local as feeling “more like regular Codex re-skinned to be less intimidating to non-software-developers,” and then sets it aside. His analysis, and this piece, focus on Work Cloud.

Access is tiered: both flavors of Work require the $20/month-and-up plans. Free users and $8/month Go subscribers are excluded. Work sessions are billed against the Codex allowance, while regular Chat sessions draw from a separate pool — a detail Willison uses to explain why the two surfaces offer different model lineups. Work exposes GPT-5.6 Sol, Luna, and Terra with reasoning levels from Light up to Ultra (Ultra, based on his Codex experience, being a mode that more eagerly delegates to sub-agents), while Chat offers its own ladder capped at High for $20 subscribers, with Pro tiers reserved for $100/month plans.

The capability list OpenAI doesn’t publish

The core of Willison’s frustration — and the reason his post exists — is that OpenAI “explain Work in terms of what it’s for, not what it actually does.” The official guidance says to use Chat “when you want an answer” and Work “when you want ChatGPT to complete a task with a clear outcome.” Willison finds this “almost entirely useless,” noting he has used regular ChatGPT for all of those task categories for years. So he built the feature delta himself. What Work has that Chat lacks:

Code execution with internet access. This is the one that made Willison, a long-time champion of the Code Interpreter pattern OpenAI pioneered in 2023, sit up. Historically, ChatGPT’s Python sandbox has been sealed behind a container proxy — no installing packages, no talking to APIs. Work’s environment can be configured with an allowlist of domains, “but the default appears to be open to all.” The practical consequence: you can have Work clone a GitHub repository, install its dependencies, and then use those tools to interact with the rest of the web. He notes Claude’s competing analysis container has allowed restricted internet access since September 2025, but only via a very short domain allowlist (PyPI, npm, GitHub clones) — Work’s default openness goes well beyond Anthropic’s posture.

A full headless Chrome browser. Work can launch a genuine Chrome instance, load websites, fill out forms, and take screenshots. When a site requires sign-in, the browser can hand control to the user to enter passwords and 2FA codes without those credentials round-tripping through the model. It can also execute arbitrary JavaScript against the DOM of loaded pages — Willison demonstrates with a snippet using Playwright’s evaluate() to extract every heading from his own site, and remarks that it feels like his shot-scraper javascript tool, “only now I can access it on my phone!”

A persistent, shared filesystem. Chat sessions get an ephemeral filesystem that dies with the conversation. Work mounts a /workspace volume across sessions: each session gets a scratch folder (his has accumulated 171), files persist between chats, and — critically — the volume appears mounted in all concurrently running Work sessions, so file edits from one are instantly visible to another. Processes and localhost servers do not cross session boundaries, but data does.

ChatGPT Sites. Work can build and deploy entire websites to Cloudflare Workers — HTML, JavaScript, server-side features, and stateful backends on Cloudflare D1 and R2. Willison’s demo prompt, delivered in one shot: “Figure out all of the places in London with a pelican in her piety, then turn that into a JSON file and build a ChatGPT sites site about them.” The result is a live, publicly reachable census of medieval pelican iconography across Greater London, complete with data downloads. Sites default to private but can be made public or shared with specific teammates.

Sub-agents and scheduled automations. Work can spawn parallel sub-agent sessions running Sol, Luna, or Terra — a power-user feature Chat lacks entirely. Scheduled prompt automations (e.g., “run a search to see if Waymo have announced a launch date for Half Moon Bay every day at 8am”) turned out to work in Chat as well, but in Work they compose with the other capabilities: a scheduled task can regenerate and redeploy a ChatGPT Site hourly.

Why the trifecta framing matters

Willison coined the lethal trifecta model in June 2025 to describe why AI agents are structurally hard to secure. Any system that combines access to private data, exposure to untrusted content, and a way to communicate information outward is vulnerable to prompt injection: an attacker who can plant instructions in content the agent will read — a webpage, an email, a PDF, a repository README — can attempt to turn the agent against its user. No reliable general defense exists. The industry’s mitigation playbook is isolation, allowlists, and human confirmation on consequential actions.

ChatGPT Work, as Willison maps it, checks all three boxes at once. The persistent filesystem holds whatever private material your sessions have produced. The open-by-default code sandbox and headless Chrome ingest untrusted web content constantly. And the outward channels — arbitrary network egress, JavaScript execution, public site deployment — are first-class product features. His conclusion is measured but pointed: “I’d love to hear more from OpenAI about how they protect ChatGPT Work sessions against prompt injection attacks.” He expects the answer is the same auto-review mechanism documented for Codex, but OpenAI has not said so, and he is asking, not asserting.

The timing gives the question weight. This is the same product family in which OpenAI’s internal evaluations recently found its Astra model’s agentic-coding gains significant enough to cite when winding down model supply to Cursor. The capability curve is steep, and the transparency curve — per Willison’s second critique — is not keeping up: OpenAI “still insist on hiding their system prompts and tools descriptions.” His closing line doubles as the post’s thesis: “If the ChatGPT Work documentation included the exact system prompt and tool descriptions used by the agent I wouldn’t have needed to write this post.”

What to do with this

For security teams, the action items are concrete. Treat a ChatGPT Work subscription as an agent runtime with internet egress and deployment rights, not a chatbot — inventory it accordingly. Assume anything written into /workspace is readable by every concurrent Work session, and by any injected instruction that succeeds in commanding one. Prefer configuring the domain allowlist over the open default if your plan exposes the control. And watch for OpenAI’s answer to the injection question: until auto-review or equivalent protections are documented for Work specifically, the trifecta framing stands as the most accurate threat model available for the product millions of subscribers now have switched on by default.