← All posts / Research

1,000 Repos, 5,000 Skills, 134% More Medals: BAAI's DisCo Turns GitHub Into Agent Food

BAAI's DisCo framework distills 1,000 widely used ML repositories into 5,000+ verified agent skills at roughly $40 per repo — and the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, with GPT-5.5 held fixed.

1,000 Repos, 5,000 Skills, 134% More Medals: BAAI's DisCo Turns GitHub Into Agent Food

The most important knowledge in machine learning is not written down anywhere an agent can use it. It lives in README files, in issue threads, in the particular order of flags someone discovered after three failed runs, in the recovery steps that rescue a training job from a corrupted checkpoint. BAAI’s new DisCo framework, published this week as arXiv:2609.02749 (“Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills”), is a serious attempt to fix that — and its headline numbers are hard to ignore: a research agent equipped with distilled skills scored 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the identical agent running without them.

The missing layer: operational knowledge

The paper’s framing is precise. Autonomous ML research agents today combine a model backbone with a harness for planning, execution, memory, and verification. What that architecture leaves out is what the authors call operational knowledge — “the know-how that separates knowing a method from making it work.”

That knowledge is not absent from the field. It is scattered across repositories and papers, but “in forms written for human readers and too large to load during a task.” An agent facing a vLLM deployment question cannot ingest a 40,000-line codebase mid-task, and even if it could, the thing it needs — when this capability applies, what to run, how to validate the result, how to recover when an experiment fails — is not a summary of the code. It is a procedure.

DisCo’s answer is to distill that procedure once, verify it, and reuse it. “Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run.”

Two distillation modes, one library

DisCo runs in two complementary forms:

  • Task-agnostic distillation condenses the field’s widely used repositories into reusable skills, applied across the open ecosystem. This is the bulk-production mode that generated the library.
  • Task-oriented distillation produces the skills a concrete task calls for, on demand.

The task-agnostic track, applied across the open ecosystem, produced the AREX-Skill Library: 5,000+ verified skills distilled from 1,000 widely used ML repositories, organized into 20 research areas and 178 capability (package) families. Coverage spans ML engineering, LLMs, computer vision, data science, scientific computing, model deployment, training infrastructure, robotics, generative media, and biomedical AI — with named skills for vLLM, SGLang, AlphaFold2, FAISS, Unsloth, Diffusers, and LeRobot, among others.

Critically, this was not free. The team reports a cost of roughly $40 per repository distilled — meaning the entire 1,000-repo library represents on the order of $40,000 in distillation compute, a strikingly small figure for an asset that more than doubles a strong agent’s medal rate on Kaggle-style tasks.

The numbers that matter

All benchmark comparisons hold the agent setup fixed — GPT-5.5 backbone, research harness, and downstream execution budget — and change only whether the agent has AREX-distilled skills:

BenchmarkScenarioMetricBaseline+ SkillsGain
MLE-benchML engineering (75 Kaggle competitions)Any Medal rate31.11%72.89%+134.3%
PaperBenchPaper replication (20 papers)Replication score29.4539.59+34.4%
FrontierCSAlgorithm optimization (188 tasks)Score70.6377.14+9.2%
PassNetCompiler pass optimization (200 samples)AS Score1.3431.531+14.0%

The MLE-bench result is the standout. Moving from a 31% to a 73% any-medal rate across 75 Kaggle competitions — without changing the model, the harness, or the compute budget — is a larger effect than most frontier model upgrades deliver in a generation. And it comes from adding context, not parameters.

The team’s own reading of where the gains come from is worth quoting in spirit: operating knowledge matters; the advantage is strongest on difficult tasks, where skills help the agent avoid expensive unguided trial-and-error and recover from near-failure states; and the budget is spent more productively, because a relevant skill graph helps the agent reach a useful region of the solution space earlier and spend more of its budget on experiments and validation.

How a skill is built — and kept honest

DisCo Creator constructs each skill through a four-stage workflow: scope capabilities from an anchor, ground them in admissible evidence, construct a candidate skill graph, and verify and refine it before publication. The anchor can be a source (task-agnostic) or a problem (task-oriented). Supporting evidence, validation checks, and unresolved gaps are retained in the construction record — provenance is a first-class citizen, not an afterthought.

Each skill uses the open Agent Skills format (a SKILL.md core with optional references/ and scripts/ resources) and is packaged so a router can narrow a request to an area, family, repository, and workflow, letting the agent load only the branch it needs. This progressive-disclosure design keeps initial context small while preserving access to deeper instructions and executable helpers.

Maintenance is addressed head-on: repository skills are tracked against upstream changes, with source commits, validation steps, and refresh requirements documented per skill. Given how fast ML tooling moves, this refresh discipline — not the initial distillation — will decide whether the library stays valuable in a year.

Try it today

The stack is fully open and usable now:

# Install the DisCo CLI (requires Node.js >= 22.19.0)
npm install -g --ignore-scripts @arex-skill/disco

# Install the skill library and start Researcher mode
disco repo-skills install
disco

Skills can also be exported into other coding agents — Codex and Claude Code are the documented targets — so the library is not locked to DisCo’s own runtime. The repository is Apache 2.0, with the important caveat that every skill carries its own license, derived from its upstream repository; users must check the per-skill license metadata before redistribution.

Why this matters beyond one paper

There is a lesson here that extends well past BAAI. For two years the industry’s implicit theory of agent improvement has been “more parameters, more reasoning tokens.” DisCo’s results argue that a large fraction of what we call agent capability is actually missing context — and that context can be manufactured, verified, and shipped as an artifact.

It also reframes the value of open source itself. The AREX-Skill Library only exists because thousands of engineers published working, documented ML repositories. BAAI’s acknowledgement section makes the point explicitly: the repo skills “exist because many researchers and engineers have released high-quality ML, agent, data, bio/chem, vision, and infrastructure projects for the community to build on.” GitHub was always a knowledge bank; DisCo is the machinery that turns deposits into agent-consumable capital.

The open questions are real. Skill quality depends on distillation-model quality, so errors in the distillery propagate into thousands of downstream runs. The per-skill licensing situation is genuinely complex — 1,000 upstream repos carry dozens of incompatible licenses. And the +134% MLE-bench gain is measured against a skills-free baseline; as skill libraries become standard equipment, the marginal advantage will compress toward the new normal.

But as a proof of what operational knowledge is worth when you package it properly, arXiv:2609.02749 is one of the cleanest demonstrations of the year. The gap between an agent that knows a method and an agent that can make it work turned out to be about $40 per repository.