One API Call, One Agent: OpenAI Puts the Codex Harness Behind the New Agents API
OpenAI's Agents API public beta exposes the managed Codex harness — sessions, compaction, subagents and sandboxes included — to every developer. No extra fee, but US-only data and no ZDR for now.
For the past two years, if you wanted an agent that could run for hours — investigating incidents, reviewing documents, reproducing GitHub bugs — you had two unattractive options: glue together your own orchestration loop on top of the Responses API, or adopt a framework like the open-source Agents SDK and shoulder the plumbing yourself. On September 10, 2026, OpenAI collapsed that stack into a single product. The Agents API, now in public beta for all developers, exposes the same managed Codex harness that powers OpenAI’s own coding agent — sessions, orchestration, context compaction, and recovery — behind one API call.
What actually shipped
The Agents API is a managed runtime built on the open-source Codex harness. OpenAI hosts and maintains the harness; your application supplies the tools and picks the execution environment. The docs organize the product around four concepts:
- Agent — the model, instructions, tools, and MCP servers available to it.
- Environment — an optional sandbox where the agent accesses files, loads skills, and runs commands.
- Session — a durable agent instance that works on tasks and responds to input.
- Events and items — the inputs sent to the agent and the output it produces during a session.
A session runs in four steps: create it with configuration, hand it a task, follow progress through streaming or webhooks, then continue with a new task or steer the current turn mid-flight. The announcement demo builds an incident-investigation agent in a single client.beta.agents.sessions.create() call — model, MCP tools, multi-agent config, and sandbox all declared inline before the agent starts root-causing an elevated 5xx rate.
OpenAI’s framing is candid about where this came from: scaling Codex and ChatGPT for Work showed what long-running agents actually need. They need a harness that manages context, uses tools efficiently, and coordinates subagents. And they need infrastructure that keeps them running reliably for days — not request/response cycles measured in seconds.
The environment question
Where the agent runs is the central architectural decision. The API supports three sandbox modes — plus running with no sandbox at all:
- OpenAI-hosted sandbox — the same sandboxing infrastructure behind Codex and ChatGPT, configurable with files, packages, skills, and plugins.
- Self-hosted — you run
codex exec-serverinside your own environment; it registers with a restricted key and connects over WebSocket, with all connections outbound. - Partner sandboxes — first-class integrations with Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel.
That partner list is the quiet strategic move here. Rather than forcing every workload onto OpenAI’s infrastructure, the company is positioning itself as the orchestration layer for agents wherever they compute — a bid to become the control plane for the emerging agent ecosystem rather than just one execution venue within it.
What the harness handles for you
The features OpenAI bundles in are precisely the unglamorous engineering that separates toy agents from production ones:
- Automatic context compaction — as a session nears its context limit, earlier turns are summarized away. Developers no longer write their own compaction logic, historically one of the more brittle parts of any long-running agent.
- Tool search — tool definitions load only when needed, reducing token usage and cost while preserving the model’s cache.
- Programmatic tool calling — agents run calls in parallel and chain operations in code, filtering and combining results so only relevant data re-enters context.
- Subagents — with multi-agent support enabled, the main agent splits complex tasks into independent pieces; each subagent keeps its own context, and the coordinator combines results.
- Session resumption — agents pick up where they left off, durable across turns without rebuilding conversation state.
OpenAI maintains the harness alongside its models, with versioned access at each model launch — meaning the runtime and the frontier models it orchestrates evolve together rather than drifting apart.
How it stacks up
OpenAI’s own comparison positions three runtimes on a spectrum. The Agents SDK runs inside your application with medium integration effort and state living in your storage. The Responses API leaves execution entirely in your hands. The Agents API sits at the low-integration end: OpenAI runs the harness, saves session configuration, turns, and items, and manages execution — at the cost of handing the runtime to a vendor.
Early customer numbers are vendor-supplied and should be read accordingly: Ciridae reports evaluation scores rising from 0.71 to 0.85 with 4x lower latency on subagent flows; SafetyKit cites 60% lower cost per case after migrating its case-review workflow; Hypha claims 86% fewer failed agent responses once the harness was separated from the sandbox; Nash.ai runs thousands of long-running agents across global logistics networks on the platform.
The fine print
Two constraints matter for regulated workloads: data residency is currently US-only, and Zero Data Retention is unsupported — and notably, choosing a self-hosted sandbox does not make the API ZDR-eligible, since session state still lives with OpenAI. Enterprises with EU data-residency requirements or strict retention policies will need to wait.
Pricing is usage-based with no extra platform fee: model tokens at standard API rates, OpenAI tools at their standard rates, and OpenAI-hosted sandboxes at standard container rates. The economics scale with what your agent actually does, not with seats or subscriptions.
Why it matters
The Agents API is best understood as infrastructure consolidation in the agent layer. Every team building long-running agents has been re-solving the same problems — context management, tool routing, subagent coordination, sandbox lifecycle — with varying quality. OpenAI’s bet is that this layer is a natural monopoly: that the company running the frontier models should also run the harness those models think inside.
The move also raises the stakes for the frameworks and platforms that built businesses on exactly this plumbing. When the model vendor ships a managed harness with subagents, compaction, and nine sandbox partners out of the box, “we orchestrate OpenAI models for you” stops being a differentiation and starts being a feature the vendor gives away for free. The counter-argument — lock-in, US-only data, no ZDR, and the risk of building atop a beta — is real, and multi-cloud teams will weigh it. But the direction is unmistakable: the agent runtime is becoming part of the model API itself, and the distance from “idea” to “working autonomous agent” just shrank to one API call.
Sources
- [1] https://openai.com/index/introducing-the-agents-api/
- [2] https://developers.openai.com/api/docs/guides/agents-api/overview
- [3] https://www.marktechpost.com/2026/09/10/openai-launches-the-agents-api-in-public-beta-putting-the-codex-harness-behind-one-api-call/
- [4] https://x.com/OpenAIDevs/status/2098130570048045453