← All posts / Research

Turning Code Into Curriculum: Xiaomi and HKU's CodeMidas Builds 5,545 RL Environments From Source Alone

A new paper from Xiaomi's MiMo team and HKU shows that plain source code — no issues, no commits, no docs — is enough to auto-build thousands of verifiable RL tasks, lifting MiMo-V2.5 by double digits on five coding benchmarks.

Turning Code Into Curriculum: Xiaomi and HKU's CodeMidas Builds 5,545 RL Environments From Source Alone

Every lab training coding agents with reinforcement learning runs into the same bottleneck: RL needs thousands of diverse tasks, and each task needs a reliable verifier that can tell a correct solution from a plausible-looking wrong one. The standard fix has been to mine development artifacts — GitHub issues, pull requests, commits — because they come with both a natural task statement and a diff that defines ground truth. The problem is coverage: you can only build tasks where developers happened to leave a paper trail.

A paper published September 18 by researchers from Xiaomi’s LLM Core team, Peking University, the University of Hong Kong, and Renmin University of China proposes a more radical starting point. CodeMidas (arXiv:2609.22068) constructs executable RL environments using source code as its only task-specific input — no issues, no commit history, no documentation required. The name is a nod to what the system does: like a Midas touch, it converts ordinary existing code into training gold.

Why existing approaches hit a ceiling

The paper’s related-work section maps the field neatly. SWE-bench-style pipelines derive task statements directly from issues and pull requests. R2E-Gym generates tasks from commits. SWE-smith synthesizes bugs by mutating code until existing tests break. MindForge uses documentation plus reference programs. All of these tie task creation to the coverage of recorded changes, tests, or documentation — which means vast amounts of implemented, working functionality in open-source codebases are simply invisible as training material.

CodeMidas flips the logic. Implemented functionality already contains both the basis for a task and a candidate solution: public interfaces and observable behavior define what an agent should build, while executing the original code provides evidence for test expectations. The surrounding codebase can be adapted into a realistic development starting point that preserves real project structure and dependencies.

An agentic pipeline, built by agents

What makes the pipeline distinctive is that agentic compute is allocated to every stage of environment construction:

  1. Task design and codebase adaptation — agents explore implemented functionality in a repository, formulate behavioral task statements around it, and adapt the codebase into a starting state where the target functionality still needs to be implemented. The task statement is deliberately behavioral: it must make required behavior explicit while leaving internal implementation choices open, so alternative correct solutions aren’t rejected.
  2. Execution-grounded test construction — agents write tests informed by actually executing the original code, with consistency checks to make sure the tests encode real behavior rather than guesses.
  3. Environment preparation — each task ships with a runnable development environment; candidate tasks must pass execution checks in fresh containers. The validation protocol is strict: two runs from the starting codebase must fail, and four runs with the reference solution in place must pass, screening out flaky environments and verifying the expected fail-to-pass transition.
  4. Post-rollout filtering — adversarial rollouts probe for exploitable leakage (ways to pass tests without doing the real work), solution reviews check verifier decisions against stated requirements, and rollout success rates guide final task selection.

The result: 5,545 verifiable training tasks extracted from 3,185 open-source codebases, spanning 23 programming languages and 15 technical domains. Python leads at 21.4% of tasks, followed by TypeScript (18.3%), Go (16.2%), C++ (12.5%), and JavaScript (11.3%). On the domain side, systems software (17.4%), web technologies (14.6%), and developer tools (13.6%) together account for nearly half the dataset.

The numbers

The team trained Xiaomi’s MiMo-V2.5 model on the full task set using GRPO with binary execution rewards — no reward model, no learned verifier. The results improved across all five external benchmarks:

  • DeepSWE (issue repair): pass rate up from 10.0% to 21.7% (+11.7 points)
  • ProgramBench (whole-program construction): Almost Solved score up from 4.5 to 21.5 (+17 points)
  • Terminal-Bench v2.1 (terminal work): pass rate up from 63.7% to 72.2% (+8.5 points)
  • Gains also held on SWE-bench Pro and RepoZero C2Rust (code translation), completing a sweep across repository repair, code translation, program construction, and terminal work.

That transfer is the key finding: tasks built from existing functionality generalize across forms of software work the training tasks never directly resembled. On the held-out CodeMidas Val set (200 tasks disjoint from training), pass rate rose from 35.0% to 44.7% and stayed 8–10 points above the initial policy from training step 40 onward.

Quality beats quantity

One of the most practically useful ablations compares task scale against task quality. Training on 1k → 3k → 5,545 filtered high-quality tasks yields progressively higher scores on SWE-bench Pro, DeepSWE, and CodeMidas Val. More strikingly, a 3k filtered subset outperforms an 8k unfiltered baseline — roughly 8,000 tasks sampled before environment cleaning, execution-consistency checks, and post-rollout filtering — on all three evaluations. The full filtered dataset beats the vanilla 8k sample by 0.59, 4.59, and 4.49 percentage points respectively. In coding RL environment construction, the filtering pipeline isn’t overhead; it’s the product.

The agents learned how to work

Beyond scores, the trajectory analysis is quietly fascinating. Comparing early and late training rollouts, the RL-trained agent doesn’t just get more answers right — it works differently:

  • Codebase exploration (read/search calls before the first edit) rose from 27.2 to 40.1 calls
  • Code drafting ratio (how much of the code it writes was already reasoned through beforehand) rose from 0.358 to 0.629
  • Distinct post-edit self-verification commands rose from 2.03 to 2.53

The agent reads more before touching anything, plans code more deliberately, and checks its own work in more varied ways — and these behavioral shifts also appear on held-out external tasks, providing evidence of genuine generalization rather than benchmark overfitting.

Why it matters

The timing is notable: this lands amid an industry-wide “agents” surge (AI Weekly’s entity tracker shows agent-related coverage up over 900% this week) and a parallel explosion of interest in coding-RL infrastructure. CodeMidas’s claim — that the hundreds of millions of lines of already-written open-source code constitute a largely untapped, self-verifying curriculum — is economically significant because it decouples task supply from the availability of development artifacts. Any codebase, however sparsely documented, becomes potential training data.

There are caveats. Building environments is itself agentic compute, so the pipeline trades human labeling cost for inference cost. The behavioral specs are model-written, so verifier quality is bounded by the constructing agent’s own understanding. And the fail-to-pass validation protocol, while rigorous, means tasks where the original code behaves inconsistently get dropped — a sensible filter, but one that shapes the data distribution.

Still, the direction is clear. As the paper puts it, the results “establish source code as a scalable foundation for constructing RL environments.” When the next generation of coding agents gets noticeably better at navigating unfamiliar repositories, curricula like this — mined not from what developers wrote about their code, but from the code itself — will be part of the reason.