← All posts / Models

Four Times Faster Than the Excel Champion: OpenAI Pitches GPT-6 Astra as the Work Model

OpenAI's new 'GPT-6 Astra for work' post puts hard numbers on the business pitch: 57.9% on Terminal-Bench 4.0, 89% fewer unintended outcomes in safety tests, Excel competition tasks at 4x human speed, and new enterprise controls and plugins.

Four Times Faster Than the Excel Champion: OpenAI Pitches GPT-6 Astra as the Work Model

When OpenAI launched GPT-6 Astra on September 3, the conversation was dominated by the “AGI era” framing, the FrontierMath Tier 4 results, and a demand surge so strong the company paused new ChatGPT Pro sign-ups. Eight days later, the company has published a follow-up — “GPT-6 Astra: The next generation in intelligence for work” — that swaps the philosophy for a spreadsheet. The new post is the business case: benchmarks priced per task, safety measured in unintended outcomes, and an Excel world-championship speed run that borders on the absurd.

The message is aimed squarely at buyers, not benchmark watchers. And it arrives at a moment when the frontier-for-work market is a knife fight: Anthropic’s Claude Fable 5.1 took the Artificial Analysis Intelligence Index crown two days after Astra shipped, Microsoft and Amazon are racing each other to host the model in their enterprise clouds, and OpenAI itself has been telling staff it may need to pace its fastest models. The “for work” post is OpenAI’s answer to the question every procurement team is now asking: which frontier model actually delivers finished work at a defensible cost?

Four Times Faster Than a World-Class Human

The post’s most quotable claim comes from esports — financial-modeling esports. GPT-6 Astra, operating through computer use, can complete Financial Modeling World Cup challenges about four times as fast as the winning human competitor. That reference is the 2023 Microsoft Excel World Championship, the sport where top analysts build full financial models against the clock. Astra isn’t just answering questions about spreadsheets; it is driving the application itself, reading the screen, planning the model, and executing keystrokes end-to-end.

That framing matters because it is the entire product thesis in miniature. Most AI systems, OpenAI argues, force businesses to prepare their data, redesign workflows, and build custom integrations before delivering any value. Astra’s pitch is inversion: it works through the same applications people already use — even when those applications have no API at all. No data plumbing project. No integration sprint. The model opens Excel, the browser, and the legacy CRM just like a new hire would, and learns the job by doing it.

OpenAI says early customers are already testing that promise in production-shaped settings: optimizing GPU allocations, spotting discrepancies in financial statements, and producing on-brand slide decks. The company also turned the model on itself — Astra was rolled out internally weeks before launch, and OpenAI’s engineering team used it to uncover a memory-allocation bottleneck that was slowing Codex sessions in a test environment. Switching allocators on Astra’s recommendation produced 25× lower turn latency at roughly 30% higher peak memory use — a systems-tuning win found by the model, not by a profiler.

The Numbers That Matter to a Budget Owner

For workloads that do run through the API, the new post anchors the efficiency argument in Terminal-Bench 4.0, which tests agents on complex terminal-based tasks spanning software engineering, system configuration, and data analysis. GPT-6 Astra reaches 57.9% — a new high for the benchmark — against 55.8% for Claude Fable 5.1 and 37.3% for OpenAI’s own GPT-5.6 Sol. The more consequential figure is cost: OpenAI estimates Astra completes tasks at approximately 9% lower API cost per task than Fable 5.1 — and 63% lower than GPT-5.6 Sol. Beating your own predecessor on price-performance by two-thirds in a single generation is the kind of arithmetic that moves enterprise budgets.

The pricing floor is unchanged: $10 per million input tokens and $50 per million output tokens, with Azure listing cached input at $1 and long-context rates at $20/$75. But OpenAI’s claim is broader than the sticker price — Astra has been trained to complete tasks in fewer tokens with fewer retries, which means less rework and a lower true cost per finished task. On the cost-efficiency frontier for professional work and coding evaluations, including Terminal-Bench 4.0 and the Artificial Analysis Intelligence Index, OpenAI says it now occupies the majority of positions.

Safety Now Has a Number: 89% Fewer Unintended Outcomes

The most technically significant part of the post may be the safety quantification. Giving an agent access to business systems is a governance problem before it is a capability problem, and OpenAI has now put a figure on alignment-in-practice: on its internal computer-use safety benchmark — which tests the hardest business scenarios, like exposing confidential information, sharing a dashboard too broadly, or deleting data — GPT-6 Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1. Additional confirmation steps and automated review improved the numbers further.

That benchmark arrives alongside new enterprise admin controls: organizations can restrict Astra’s access to approved websites and desktop applications, manage uploads and downloads, and control browsing history. Confirmation policies can require human approval before consequential actions, and automated review inspects potentially unsafe or unauthorized tool calls. The design philosophy is explicit — start with a limited configuration and expand access over time, rather than trusting a frontier model with everything on day one.

This is also the first model to reach the “Critical” cybersecurity capability threshold under OpenAI’s Preparedness Framework — capable enough at finding and exploiting zero-day vulnerabilities that the company treats it as a dual-use risk. OpenAI says it has strengthened protections against both misuse and unauthorized model actions, including resistance training against attempts to bypass safeguards and automated checks designed to block harmful responses. Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval.

The Distribution Play: Desktop Plugins and Two Clouds

Two distribution moves complete the business pitch. First, OpenAI is launching new enterprise plugins in ChatGPT Desktop powered by the latest browser-use capabilities — from Oracle Analytics, Power BI (a Microsoft Fabric service), Navan, and Avalara. The plugin list is a tell: analytics, expense management, and tax compliance are precisely the back-office surfaces where an agent that can log in and drive a UI beats one that waits for an API.

Second, the cloud majors are treating Astra as a land-grab product. Microsoft made GPT-6 Astra generally available in Foundry on day one, with Standard and Provisioned Throughput deployment options in Global and US Data Zone geographies — the Azure post emphasizes scoped credentials, human checkpoints for consequential actions, and activity records aligned to organizational risk requirements, with prompts and outputs never used for training. Amazon, not to be outflanked, brought the model to Bedrock the same week. When your frontier model is simultaneously a first-party OpenAI product, an Azure Foundry SKU, and a Bedrock endpoint, the marginal enterprise buyer no longer has a vendor-lock excuse to delay adoption.

The Sober Read

Three caveats deserve equal billing. First, OpenAI’s own comparisons come from OpenAI’s evaluations — the 4x Excel figure, the 89% safety improvement, and the cost-per-task estimates are all self-reported, and rivals publish rival numbers. Second, independent leaderboards tell a more muddled story than the blog post: Artificial Analysis’s v4.2 index still ranks Claude Fable 5.1 first, with Astra configurations second and third, and the two labs are trading category leads benchmark by benchmark rather than one model dominating. Third, the demand that made OpenAI pause Pro sign-ups has not fully subsided — a model that is 63% cheaper per task is only cheap if there is capacity to serve it.

But the direction is unmistakable. The launch-week question — “is this AGI?” — has already been replaced in enterprise procurement decks by a different one: “what does a finished task cost?” OpenAI’s answer, published three hours before this writing, is that the world’s most capable model is also the one that closes tickets, builds models, and drafts decks at the lowest cost per unit of work — while producing 89% fewer of the unintended outcomes that keep CISOs awake. That is not an AGI claim. It is a line item. And in the work-software market, line items win.