← All posts / Tools

The 47x Premium Question: CodeRabbit's Independent GPT-6 Astra Code-Review Verdict

CodeRabbit's independent evaluation finds GPT-6 Astra catches 22% more bugs than Claude Opus 5 in code review — with gains jumping to 33% on hard cross-file reviews, but at 2.5x Sol's cost and 47x Luna's.

The 47x Premium Question: CodeRabbit's Independent GPT-6 Astra Code-Review Verdict

Two days after OpenAI shipped GPT-6 Astra, the first rigorous third-party verdict on what the model actually does to real engineering work has landed — and it comes from an unusual referee. CodeRabbit, the AI code-review platform that processes pull requests for thousands of teams, published its early evaluation of Astra on September 4, and the numbers tell a more nuanced story than either the launch benchmarks or the skeptical hot takes that followed them.

The headline finding: in CodeRabbit’s evaluation, Astra caught approximately 4% more labeled bugs through actionable findings than GPT-5.6 Sol, OpenAI’s own previous flagship — and 22% more than Anthropic’s Claude Opus 5. But the aggregate number hides where the real movement is. On the harder cross-file subset, where a change looks correct in isolation but breaks code elsewhere in the system, Astra’s relative advantage grows to 20% over Sol and 33% over Opus 5.

Why cross-file reasoning is the whole ballgame

Some of the hardest work in code review happens outside the changed lines. A diff can pass every local check and still break a distant module that depends on an implicit assumption the author never knew existed. This is precisely the class of bug that human senior engineers catch and junior reviewers miss — and precisely where CodeRabbit’s data says Astra pulls ahead.

The company’s interpretation is carefully hedged: Astra’s most interesting advance lies in connecting the right information. A large context window creates room for information, but useful reasoning requires the model to identify which pieces matter, link them, and reach a conclusion supported by evidence. The larger gains on scattered-evidence tasks suggest real progress on that connective work — though the evaluation “does not isolate the cause of that progress or prove that more context alone improves a model’s performance.”

It is an early, directional result, and CodeRabbit says so explicitly. These numbers describe one part of review performance; they do not establish an overall ranking of review quality, predict a team’s defect rate, or promise the same gain on every pull request. The overall 4% gain over Sol looks modest partly because the evaluation includes simpler reviews, where a stronger model has less room to differentiate itself.

The pricing reality check

Then comes the bill. Astra’s standard API rates are $10 per million input tokens and $50 per million output tokens — identical base rates to Anthropic’s Claude Fable 5.1, though caching prices differ. CodeRabbit puts those rates in context with an illustrative task using 100,000 uncached input tokens and 10,000 billable output tokens:

ModelInput / 1MOutput / 1MIllustrative task cost
GPT-5.6 Luna$0.20$1.20$0.032
GPT-5.6 Terra$2.00$12.00$0.32
GPT-5.6 Sol$4.00$20.00$0.60
GPT-6 Astra$10.00$50.00$1.50
Claude Fable 5.1$10.00$50.00$1.50

At fixed token usage, Astra costs 2.5 times Sol, about 4.7 times Terra, and about 47 times Luna. Those are meaningful premiums, and CodeRabbit is careful not to overstate them: they are not predictions of cost per completed task. A model that needs fewer tokens or fewer attempts could narrow the gap — and OpenAI reports lower estimated task costs for Astra in some of its own evaluations despite the higher token prices. The company’s advice is to measure total cost per successful outcome on your own work rather than assuming either token price or a capability score settles the decision.

The practical guidance that falls out of this: route the hard, scattered-evidence work to Astra and let cheaper models handle the routine checks. “That doesn’t mean switching every session to Astra, maxing out reasoning effort, and letting it rip.”

NIGHTSHIFT: the strangest benchmark of the year

The evaluation’s most surreal section is the game. CodeRabbit used Astra to build NIGHTSHIFT, a complete action RPG written in Godot and GDScript — not as a toy demo, but as a stress test of exactly the capability the code-review numbers highlight: reasoning over how one change ripples through an interconnected system.

The game spans seven character classes, a 988-node passive skill tree inspired by Path of Exile’s legendary node tree, active skills, runes and socketable upgrades, skill evolutions, and co-op play — across 40 zones in 10 acts. The hardest problem, the team writes, was balancing the interactions between systems, then revisiting that balance as the game changed. Changing one class alters which upgrades are useful, how a skill develops, and what a party can handle.

CodeRabbit made fundamental changes to core systems mid-development and asked Astra to work through the consequences and rebalance. The model also built native PS5 and Xbox controller support, native macOS, web, and Linux builds, and a LAN/same-machine co-op mode that required generating a new Xcode project, an App Store Connect account, multiple certificates and entitlements, and notarization — “autonomously, only pausing to occasionally ask for authority and permissions it didn’t already have.” As a bonus, the game was built so agents themselves could play it: developers found themselves in live co-op sessions with Astra as a teammate. “I’m working on model evals,” one engineer joked, “became a useful explanation for having a game open when a manager stopped by.”

The privacy footnote that matters for enterprise

For a code-review vendor, model choice is also a data-governance decision, and the post closes with a disclosure that reflects the new reality of frontier-model deployment. CodeRabbit notes that GPT-6 Astra supports zero data retention for eligible API customers under OpenAI’s data controls. Anthropic’s Fable models, by contrast, require 30-day retention by default for safety monitoring — though eligible customers can now use Fable 5 and 5.1 with ZDR while the new Enterprise Frontier Safeguards program is introduced, which keeps retained activity data in customer-controlled infrastructure.

That asymmetry — OpenAI offering ZDR cleanly, Anthropic gating it behind an enterprise program — is the kind of procurement detail that never makes launch headlines but decides which model actually runs against a company’s proprietary codebase.

What to watch

The evaluation arrives amid a noisy week for Astra. OpenAI’s launch materials show the model scoring 59.3% on its headline agentic-coding comparison versus 55.5% for Claude Opus 5 and 53.6% for GPT-5.6 Sol, and Artificial Analysis’s independent benchmarking puts Astra at roughly parity with Claude Opus 5 and Fable 5 in coding-tool contexts. Meanwhile, the model’s system card disclosed that chain-of-thought monitorability has substantially decreased — a safety trade-off that sits uneasily beside the capability gains.

CodeRabbit’s contribution is to move the conversation from benchmark scores to unit economics: a 33% improvement on the hardest review class, priced at a 2.5x premium, with a clear recommendation to measure cost per successful outcome rather than per token. The next advance the company wants to see is making this depth of reasoning dependable — consistent gains on difficult work, conclusions people can verify, and lower total cost. Until then, the pragmatic play is surgical: Astra for the cross-file mysteries, Luna for everything else.

Sources are listed in the post metadata. Evaluation figures and quotations are from CodeRabbit’s blog post by Erik Thorelli and Erfan Al-Hossami, September 4, 2026; pricing checked against OpenAI and Anthropic published rates.