← All posts / Research

36.5% and Nowhere Left to Hide: Φ-Bench Asks If Frontier LLMs Can Engineer the Infrastructure That Powers Them

A new 85-task benchmark from USTC, StepFun and Yale grounds LLM agents in real GPU kernel, training and serving codebases — the best model, Claude Opus 5, scores just 36.53%, and on hardware-edge tasks the field collapses to 5.4%.

36.5% and Nowhere Left to Hide: Φ-Bench Asks If Frontier LLMs Can Engineer the Infrastructure That Powers Them

Can a model that runs on a datacenter full of GPUs meaningfully improve the software stack that trains and serves models like itself? That is the disarmingly simple question behind Φ-Bench (the Frontier AI Infrastructure Benchmark), released September 9 by a 13-author team spanning the University of Science and Technology of China, StepFun, Peking University, HKUST, Yale and the University of Pennsylvania. The answer, at least for this generation of frontier models, is: not yet — and the benchmark is the first rigorous, reproducible measurement of exactly how far short they fall.

What Φ-Bench actually tests

Most LLM coding evaluations live comfortably in toy territory: complete a function, fix a bug, optimize a predefined operator against a predefined target. Φ-Bench deliberately rejects that framing. Its 85 tasks are derived from optimization problems studied in frontier systems research and grounded in real-world public code repositories, and its three task formats get progressively more open-ended:

  • Kernel Function Completion (KFC, 55 tasks) — localized, kernel-level implementation work: writing the computational primitives that everything else depends on.
  • Long-Horizon Implementation (LHI, 20 tasks) — repository-scale engineering: navigate an existing infrastructure codebase, understand it, and implement substantial changes that build and pass every test.
  • End-to-End Optimization (E2EO, 10 tasks) — the most open-ended format: given a fully editable system, a wall-clock budget and a parameter-count floor, minimize a real objective such as validation bits-per-byte on a nanoGPT training setup.

To build it, the team constructed a bottom-up taxonomy of modern LLM infrastructure from 2,260 papers and 1,852 repository artifacts, then used an agent-loop-based pipeline to automatically mine high-value engineering challenges from repos and iteratively generate test cases — a pragmatic answer to the fact that manually curated engineering tasks simply don’t exist at the scale a benchmark like this needs.

The scoring is unforgiving by design. Every attempt is correctness-gated: buildability, functional correctness, edit constraints and anti-cheating checks are evaluated before any reward is issued. Performance tasks use an AB-BA paired measurement protocol with at least five measurement pairs and logarithmic normalization against a reference solution; implementation tasks are strictly binary — full credit only if the solution builds, imports, passes every test case and respects all forbidden-edit constraints.

And because frontier models are crafty, the benchmark ships with two complementary proctoring mechanisms: a rule-based monitor that scans complete solution trajectories for prohibited behaviors (searching GitHub for the original implementation, pulling the upstream patch diff, recovering code from published Python packages), plus a dedicated proctor agent that inspects submissions for subtler hacks like hard-coding expected outputs or introducing branches that artificially inflate measured performance. Caught cheating means a zero for that attempt.

The results: one-third of the way there

The headline number is sobering. Claude Opus 5 tops the leaderboard at 36.53% overall, followed by Moonshot’s Kimi K3 at 28.12% and Qwen3.8 Max at 27.73%. GPT 5.6 Sol manages 24.51%, GLM 5.2 scores 21.92%, and Claude Sonnet 5 lands at 17.58%. Keep in mind what these agents had access to: the complete repository, the test cases, and up to 16 candidate submissions per task on an NVIDIA H20 GPU with 32 GiB of memory and eight CPU cores. Even with every advantage a human engineer could ask for, the best model clears barely a third of the maximum score.

The category-level breakdown is where the story gets interesting. Claude Opus 5 leads in five of the nine infrastructure domains (Training, I/O, Data Infrastructure, Kernel, and E2EO-driven categories), but no model is uniformly strong. Kimi K3 is the best model on Inference & Serving and System Optimization; Qwen3.7 Max unexpectedly tops Hardware & Edge; GLM 5.2 dominates System Assurance with 57.40% — the single highest category score for any model outside Claude’s Training numbers.

Hardware & Edge is the graveyard: the best score across all nine evaluated models is 5.4%, and several models — including GPT 5.6 Sol and Claude Sonnet 5 — post literal zeros. Current frontier models, the authors conclude, still lack sufficient understanding of the hardware aspects of AI infrastructure.

Persistence beats brilliance

The error-mode analysis produces one of the paper’s most counterintuitive findings: models that produce more errors score higher. Claude Opus 5, Kimi K3 and Qwen3.8 Max all log large error counts, because they keep attempting hard tasks through rounds of trial, diagnosis and correction. Weaker models like DeepSeek V4Pro and Qwen3.7-Max produce fewer errors — because they give up sooner or retreat to simpler solutions. The capacity to persist through failure, not raw first-shot accuracy, separates the leaders.

Claude Opus 5’s error composition is also distinctive: a significantly lower proportion of Python runtime errors than peers, with its errors concentrated in CUDA execution instead. It writes Python that works within the repository on the first attempt, freeing its iteration budget for the genuinely hard low-level problems.

The case study on the nanoGPT E2EO task — where models must optimize a mixture-of-experts training setup through token dispatch, load balancing and capacity-factor tuning — reads like a field guide to good vs. bad experimental methodology. Claude Opus 5 builds lightweight local validation experiments to cheaply screen hypotheses before committing a formal submission, explores broadly, then focuses. DeepSeek V4Pro applied torch.compile to expert modules without first validating checkpoint compatibility and burned a submission when the checkpoint wouldn’t load. Qwen3.7 Max ran disciplined single-variable experiments whose differences were smaller than measurement noise. Strong infrastructure engineering, it turns out, looks a lot like strong science: control your variables, respect your noise floor, attribute outcomes cautiously.

One more sobering data point: reasoning budget doesn’t reliably buy performance. Across 20 LHI problems at varying reasoning-effort levels, all three evaluated models peaked at max effort, but scores did not increase consistently with budget — GPT 5.6 Sol in particular showed notable non-monotonicity.

Why this matters

The economics of the AI industry currently rest on a single bet: that algorithmic and infrastructural efficiency gains will keep compounding. Infrastructure-level optimization is where GPU utilization is won and computational cost is lost — the difference between a training run that fits in a budget and one that doesn’t. If LLM agents could reliably do this work, the loop of “AI improving AI” becomes concrete rather than rhetorical.

Φ-Bench makes the honest measurement possible: a public leaderboard, the dataset on Hugging Face, and the full harness on GitHub. It also quietly sketches the shape of the next competitive frontier. Long-Horizon Implementation scores sit consistently below KFC scores for every model — repository-scale, open-ended engineering is where everyone struggles. And the near-total collapse on Hardware & Edge tasks maps precisely onto the skill the industry most urgently needs as accelerator supply tightens and every watt and HBM gigabyte is contested.

A model that scores a third on this benchmark is a useful junior engineer with infinite patience. A model that scores 80% would be something else entirely — and Φ-Bench is now the yardstick that will tell us when that happens.

Φ-Bench is available at faibench.org, with the dataset on Hugging Face and the evaluation harness on GitHub.