← All posts / Industry

One Billion Users, 3.1 Agent-Workdays, and a Chip of Its Own: Inside OpenAI's 'The Work Now Within Reach'

OpenAI CFO Sarah Friar's September 8 essay ties the whole machine together: 1B+ weekly users, 2.5M businesses, agents logging 3.1 workdays per researcher day, GPT-6 Astra benchmarks, 20% cheaper serving, and the Jalapeño chip — one flywheel, argued in public.

One Billion Users, 3.1 Agent-Workdays, and a Chip of Its Own: Inside OpenAI's 'The Work Now Within Reach'

On September 8, 2026, OpenAI published an essay that reads less like a product announcement and more like an investor briefing. Titled “The Work Now Within Reach” and authored by chief financial officer Sarah Friar, it is the company’s most complete public attempt to explain its own economics: why consumer scale, enterprise adoption, frontier research, and a vertically integrated compute stack are not separate bets but a single reinforcing loop.

The timing matters. OpenAI has spent 2026 absorbing enormous capital commitments — data centers, chip designs, credit structures — and the essay is effectively the argument for why that spending converts into revenue rather than into an expensive science project.

The flywheel, stated plainly

Friar’s thesis is that consumer and enterprise usage strengthen each other. Every research advance improves ChatGPT, ChatGPT Work, Codex, and the API; every new user generates usage data and demand that funds the next research cycle; and a full-stack compute strategy — from custom silicon to serving software — keeps the marginal cost of intelligence falling.

The scale figures anchor the argument: more than one billion weekly active users and 2.5 million businesses now touch OpenAI products. Free access, supported by advertising, serves as discovery; subscriptions and usage-based pricing let customers spend more as they find more value.

To back the “people find more value” claim, OpenAI studied individual ChatGPT plans: six months after signup, daily message volume ran roughly 50% higher than in the first month, and users had tried roughly twice as many distinct tasks. The methodology is disclosed — a 0.1% sample of accounts created between October 15, 2025 and May 1, 2026, tracked through May 31, 2026, with messages sorted into 53 capability categories and measured against each user’s first 28 days. It is a retention-and-expansion story told with unusual granularity.

Agents: 3.1 workdays per human workday

The essay’s most striking number is borrowed from OpenAI’s September 6 research update: across its research organization, agents now contribute 3.1 workdays of runtime for every human workday, measured against a standard eight-hour day as of mid-August 2026. Before June 2026, total agent runtime still sat below total human labor time. The crossover happened in a single summer.

The companion figures are just as aggressive: the median researcher (ranked by agent usage) was consuming more than $600 per day of inference at API prices by mid-August, and the 90th-percentile researcher more than $7,000 per day of tokens.

OpenAI is careful with the framing, and readers should be too. An agent-workday is not a human-workday of equivalent output — machines pursue dead ends and produce work that needs correcting. And success still requires steering: across the prior six months, more than half of successful tasks estimated at four to eight hours of human effort involved at least one human intervention. As The New Stack observed in its skeptical read of the same data, more agent hours do not automatically mean faster breakthroughs — supervision capacity may be the real bottleneck. Humans still set research priorities and judge results.

Five customer deployments, quantified

The enterprise half of the flywheel is illustrated with five deployments, each figure coming from the customer’s own account as presented by OpenAI:

  • Boston Children’s Hospital reported more than 40 diagnoses in previously unresolved rare-disease cases through AI-assisted research.
  • Cars24 said its agents handle more than one million conversation minutes per month across car buying and selling.
  • Circles reported 65% autonomous resolution of customer-service interactions across supported workflows.
  • Balyasny Asset Management said its Central Bank Speech Analyst cut macroeconomic scenario analysis from two days to roughly 30 minutes.
  • Replit offers a Free Mode that lets users plan software without consuming their usage allowance.

Individually these are case studies; together they sketch the demand side that Friar’s investment framework requires — OpenAI says it judges each investment by the demand it can serve, how quickly it becomes productive, and whether returns justify the capital committed.

Astra, serving costs, and the Jalapeño chip

Capability-wise, the essay frames GPT-6 Astra as the current engine: state-of-the-art (by OpenAI’s own benchmarks) in computer use, browsing, software engineering, cybersecurity, science, and professional work — 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench, meeting the Critical threshold for cybersecurity under the company’s Preparedness Framework. Astra began rolling out to a limited set of organizations on September 3, 2026, with general availability across ChatGPT paid tiers, the API, Microsoft Azure, and AWS Bedrock in the following days.

On the cost side, the essay reiterates two results from OpenAI’s July 29 GPT-5.6 engineering post: production serving improvements that cut end-to-end serving costs by 20%, and speculative decoding that raised token-generation efficiency by more than 15% — gains achieved with GPT-5.6 Sol operating inside Codex, including autonomously rewriting production kernels.

The most strategically loaded section recaps Jalapeño, OpenAI’s custom inference chip. In InferenceX tests across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5, OpenAI reports 1.5–1.9x the peak token throughput per watt of the commercial systems tested, with end-to-end latency 1.7–3.6x lower. Deployment is planned in OpenAI’s own compute infrastructure by year’s end, alongside accelerators from NVIDIA, AMD, and other partners — with second- and third-generation designs already in development. If those numbers hold in production, OpenAI’s margin structure starts to detach from the merchant-GPU market.

The read-through

Every number in the essay has an interested party behind it: benchmarks are self-reported, customer figures are customer-supplied, and the agent-usage snapshot is an internal view with no external audit. But that is precisely what makes the document interesting. It is OpenAI’s own accounting of how capability becomes cash flow — usage deepens over months, agents compound researcher effort, serving costs fall per token, and custom silicon aims to bend the cost curve further.

The unanswered question is the one critics keep raising: whether 3.1x agent-runtime translates into proportionally faster frontier progress, or whether the human attention required to steer, verify, and correct that output becomes the new constraint. Friar’s essay implicitly concedes the point — “people still set research priorities and judge results” — while betting that the constraint loosens with every model generation.

Either way, the flywheel is now formally documented by the person who has to make it pay.