← All posts / Research

Real-SWE: Coding Agents Graded on Code No Model Has Ever Seen — Fable 5.1 Leads at 38.8%

Specific Labs' new Real-SWE benchmark tests coding agents on licensed private enterprise codebases no model could have memorized. Fable 5.1 resolves 38.8% of tasks, GPT-6 Astra 33.8%, and open-weight GLM-5.3 surprises at 28.8% — but six of ten sample tasks sit below a 15% resolution rate.

Real-SWE: Coding Agents Graded on Code No Model Has Ever Seen — Fable 5.1 Leads at 38.8%

Every coding-agent benchmark shipped in the last three years shares one structural weakness: the tasks come from public GitHub repositories, which means the model under test has plausibly already seen the exact bug — and the exact pull request that fixed it — somewhere in its training data. On September 10, Specific Labs, a young YC F25 applied-AI research company, released a benchmark built to make that shortcut impossible. Real-SWE evaluates frontier coding agents exclusively on private, licensed, production enterprise codebases: systems that have never been on the public internet and that no model could have trained on. The first leaderboard went live this week, and it is sobering.

The premise: 99% of enterprise tokens are invisible

The core argument behind Real-SWE is simple. As Specific Labs puts it, “99% of tokens in real-world enterprises are hidden away from the frontier models.” Public benchmarks — SWE-bench and its many descendants — measure how well an agent recognizes patterns common across open-source Python and JavaScript projects. Real-SWE measures something different: whether an agent can walk into an unfamiliar proprietary system, read its business logic the way a new engineer does in their first week, and make a change that works with what is already there.

To build it, Specific Labs licensed production codebases from real companies with substantial usage and demanding workloads. The examples they cite include a Luma/Partiful competitor with over 200,000 users and a top-100 App Store ranking, a consumer fintech platform processing 100,000+ bank statements, and enterprise AI sales platforms supporting complex business workflows. Every task is “inspired or lifted verbatim” from work assigned to a salaried engineer — billing fixes, tax calculations, customer migrations — meaning each task has a direct relationship to spend.

The tasks are deliberately underspecified, about par with DeepSWE and Terminal-Bench: a typical instruction runs 1,742 characters, and the median reference solution touches 11 files, versus 6 for both FrontierCode and DeepSWE. Agents run in their native harnesses — Claude Code, Codex CLI, Gemini CLI, Grok Build, Kimi Code, Muse Code — because the benchmark evaluates model-and-harness combinations, the way enterprise engineers actually work. Each task was run eight independent times per model, with resolution rate equivalent to pass@1.

The leaderboard: Fable 5.1 first, GLM-5.3 the surprise

The headline results, with estimated cost per rollout:

#Model (harness)Resolution rateCost/rollout
1Fable 5.1 (Claude Code)38.8%$6.96
2GPT-6 Astra (Codex CLI)33.8%$4.67
3Gemini 3.8 Flash (Gemini CLI)31.2%$2.50
4GLM 5.3 (Claude Code)28.8%$5.12
=5Grok 4.6 (Grok Build)23.8%$3.44
=5Muse Spark 1.3 (Muse Code)23.8%$2.74
7Kimi K3 (Kimi Code)18.8%$3.90
8GPT-5.6 Sol (Codex CLI)16.2%$2.65

Two results stand out. First, Anthropic’s Fable 5.1 confirms its reputation as the strongest agentic coder, but at $6.96 per rollout it is also the most expensive — Gemini 3.8 Flash gets within 7.6 points of it at roughly a third of the price. Second, the open-weight GLM-5.3 landing in fourth at 28.8% has drawn the most community reaction. As explainx.ai noted, a strong Real-SWE showing is a different kind of evidence than Z.ai’s own benchmarks: it suggests GLM-5.3’s coding gains generalize to genuinely unfamiliar, private enterprise code — a meaningfully higher bar than most open-weight claims get held to.

The task-level breakdown is where the benchmark gets brutal. On the ten sample tasks analyzed:

  • Multi-region sweep and API keys & environments are the most tractable (67.2% and 65.6% resolution).
  • Entitlement overage lines sits at 50.0%, and customer identity migration at 40.6%.
  • Then the floor drops out: billing schedule migration at 14.1%, API token metering at 12.5%, S3 datastore measurement at 10.9%, linearizable scan at 4.7%, tax jurisdiction at 3.1% — and analytics stream reducer at a flat 0.0%. Every model, all eight rollouts, failed.

Six of the ten sample tasks have resolution rates below 15%, and no model solves every task. There are surprising inversions, too: Grok 4.6 — fifth overall — went 8-for-8 on customer identity migration, a task where Fable 5.1 managed only 3-of-8. GLM-5.3 was the only model to make progress on S3 datastore measurement. Strength profiles differ sharply even when headline scores are close.

How models fail: missed requirements, not crashes

Real-SWE classifies failures using the DeepSWE taxonomy, and the distribution says a lot about what “not ready” actually means:

  • Missed requirement (leaving out behavior the instruction requires) is the most common failure mode almost across the board — 67.2% of Grok 4.6’s failed runs, 53.8% of Kimi K3’s, 38.6% of GLM-5.3’s, 36.7% of Fable 5.1’s.
  • Unverified assumption — building on a guess about the system instead of checking it — peaks with GPT-5.6 Sol at 43.3%.
  • Integration error (right idea, wired into the surrounding system incorrectly) dominates Gemini 3.8 Flash’s failures at 49.1%.
  • Regressions are relatively rare, and wrong-file deliveries — putting the change somewhere the running app never calls — are almost exclusively a GLM-5.3 and Kimi K3 behavior.

Perhaps counterintuitively, giving models more time doesn’t help: 71.4% of rollouts under 10 minutes failed, compared with 73.4% of longer ones. The bottleneck isn’t thinking time; it’s triaging multiple systems and understanding requirements buried in existing business logic.

Why this matters

For enterprises choosing a coding agent, the implication is direct: public-repo scores overstate readiness for private codebases, and a Real-SWE-style score is a better proxy for “can this agent onboard onto our system the way a new hire would.” The honest caveat is that the results show we’re far from that reality — the best model resolves well under half of real engineering tickets.

Real-SWE also crystallizes a broader 2026 evaluation trend: as leaderboards saturate, the benchmarks that still differentiate models are the ones structurally resistant to memorization. Senior SWE-Bench did it with under-specified Slack-style instructions; Terminal-Bench did it with full terminal environments; Real-SWE does it at the data-sourcing layer, with code that cannot have leaked into training. The approaches are complementary, and future benchmarks will likely combine them.

The limitations are real and worth stating: the benchmark is days old, with no independent replication; the leaderboard is run by the lab that built it; and because the codebases are private, outside researchers can’t audit every task the way they can with SWE-bench. That privacy is the point — it’s what makes the benchmark un-memorizable — but it trades auditability for integrity, and the leaderboard’s long-term credibility will depend on rotating fresh company codebases in to resist overfitting.

Still, the first data point is a valuable one. When coding agents are finally graded on the 99% of code they’ve never seen, the frontier resolves about a third of real tickets, the best open-weight model is ten points behind the leader, and the hardest business-logic tasks remain unsolved by anyone. Benchmark contamination was flattering us. Real-SWE is the mirror.

Results and methodology: realswe.withspecific.com. Resolution rates are pass@1 averaged over eight runs per task; costs are estimated per rollout as published by Specific Labs.