28% Is the New 100%: Google's Android Bench 2.0 Grades AI on Tasks That Take Humans a Week
Google's Android Bench 2.0 replaces binary pass/fail grading with continuous scoring and week-long engineering tasks — GPT-6 Astra tops the long-horizon leaderboard at 28%.
For two years, AI coding benchmarks have shared an awkward secret: once a model clears roughly 90% on a test suite, the number stops meaning anything. Google’s answer, published September 17 on the Android Developers Blog, is to make the test radically harder. Android Bench 2.0 — a major upgrade to the company’s official benchmark for AI-assisted Android development — introduces long-horizon tasks (LHTs): engineering jobs “of great complexity that take an engineer multiple days or even a week to complete,” alongside a shift from binary pass/fail grading to continuous scoring.
The result is a leaderboard that looks brutally honest. The highest pass rate any frontier model achieves on the new long-horizon tasks is around 28%, compared with roughly 91% on the original, bite-sized tasks. That single contrast is the whole story: modern coding agents remain excellent at finishing what a human started, and still struggle to carry an ambitious engineering project from blank slate to done.
What changed in version 2.0
The first iteration of Android Bench, launched earlier in 2026, focused on incremental changes to existing repositories — bug fixes and small feature requests. That reflected both the capability of AI assistants at the time and how developers actually used them. But the work people delegate to AI has grown far more ambitious, and the benchmark needed to catch up.
Android Bench 2.0 mirrors that ambition with long-horizon tasks that include:
- Upgrading dependencies across a real project
- Adding substantial new features to an existing app
- Building apps from scratch
- Converting a cross-platform app to Android
The framework has also been aligned with the Harbor benchmark framework, and Google is introducing agentic evaluation — running models inside the coding agents from their own providers, such as GPT 5.6 Sol on Codex and Gemini 3.8 Flash on Google Antigravity. This pairing matters, Google notes, because harness design measurably affects developer outcomes: techniques like prompt caching and compact tool windowing produce real token reductions that raw model-to-model comparisons hide.
Continuous scoring instead of pass/fail
The deeper methodological shift is scoring. On multi-day engineering tasks, binary grading actively misleads. Google gives a concrete example: an agent might refactor 40 screens to Jetpack Compose, set up the database tables, and satisfy 90% of the requirements — then fail a single edge-case assertion. Binary scoring records that run as 0%, “obscuring the model’s architectural capabilities.”
Android Bench 2.0 instead reports two metrics: a pass rate (how often an agent fully solves a task) and a completion rate (how much of the task it accomplished), computed from functionality, visual fidelity, and regression avoidance, with objective penalties for deviations from evaluation instructions or structural constraints. Each model’s card view on the leaderboard breaks out pass rate, completion rate, and average costs per model and per task — a welcome nod to the fact that “which model is best” is inseparable from “what does it cost to run.”
What the long-horizon tasks revealed
The LHT dataset is more than a leaderboard; it’s a fairly sharp diagnostic of where AI assistance genuinely stands in late 2026:
New code beats old code. Across model tiers, AI writes new code better than it refactors existing code. Refactors and migrations are harder because success depends on architectural complexity rather than code volume — the model has to understand a system, not just produce one.
Deterministic transformations are a solved problem. Models are consistently strong on well-established, deterministic transformations: converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer. They apply these patterns reliably even across 125+ files and 8,000+ lines of code.
Runtime validation is the frontier’s blind spot. Models struggle when tasks require runtime validation (for example, resolving missing dependency injection graphs), involve breaking framework changes, or run into knowledge gaps with unreleased libraries — the fresh-SDK problem every mobile developer knows.
Cross-platform porting remains open. Porting a cross-platform app to Android is still an unsolved challenge: no model hits a 100% pass rate, and frontier models reach at most an 80% completion rate.
The new leaderboard
Alongside the 2.0 methodology, Google expanded the roster with Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI’s GPT-6, Anthropic’s Fable 5.1, Kimi K3, and Qwen 3.8 Max. At publication, OpenAI’s GPT-6 Astra sits at the top of the long-horizon leaderboard with a 28% pass rate — a reminder that “state of the art” now means clearing barely a quarter of tasks that a senior Android engineer would be expected to complete.
That framing is the benchmark’s real contribution. For all the conference-demo energy around autonomous coding agents, Android Bench 2.0 quantifies the gap between “impressive demo” and “delegable engineering work” — and it does so on tasks drawn from the actual daily reality of the world’s largest mobile platform, with millions of developers as its audience.
Why it matters
Benchmarks shape incentives. When SWE-bench-style suites saturate, labs optimize for narrow patches; when a major platform owner publishes a public, continuous-scored, cost-transparent leaderboard built on week-scale tasks, the optimization target moves toward exactly the capabilities developers keep asking for: architectural reasoning, regression discipline, and trustworthy long runs. Google also promises more to come — additional model-and-agent combinations “in the coming weeks,” so teams can discover which pairings actually work in practice.
The benchmark, methodology, and leaderboard are public on developer.android.com, with feedback channels on GitHub. For anyone building or buying agentic coding tools, the 2.0 numbers are the most honest map available of where the floor currently sits — and how far the ceiling still is.