← All posts / Research

1% of Tokens Can Be Enough: The Signal-to-Noise Trick That Makes Sparse Distillation Actually Work

MBZUAI and Ant Group researchers show that scoring just 0.1%–1% of tokens during on-policy distillation can match full supervision — if you select tokens by gradient reliability, not just usefulness.

1% of Tokens Can Be Enough: The Signal-to-Noise Trick That Makes Sparse Distillation Actually Work

On-policy distillation (OPD) has quietly become one of the workhorse recipes for training smaller, cheaper reasoning models: a student model generates its own rollouts, and a stronger teacher scores the next token at every position along the way. Compared with sparse verifiable rewards or sequence-level losses, that token-level signal is dense — and expensive. Every token the teacher scores costs compute, which is why a fast-growing body of work asks a deceptively simple question: do we really need to supervise every token?

A paper posted to arXiv on September 21 — “1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation” (arXiv:2609.24432), by Huanxin Sheng, Zhiling Ye, Haonan Wang, Jian Wang, Jinjie Gu, and Jian Kang, a collaboration between MBZUAI and Ant Group — argues that most of the field has been answering a slightly wrong question. Existing sparse-OPD methods rank tokens by usefulness: importance heuristics, student uncertainty, teachability, or simply “the earlier the better” (prefix-based selection). What they ignore is whether the gradient you’d actually compute at that token can be trusted when it’s estimated from a single sampled next token.

Useful is not the same as reliable

The paper’s core observation is an estimation-theoretic one. The gradient of the reverse-KL objective at a given prefix is an expectation over the model’s full vocabulary. In practice, sampled-token OPD approximates that expectation with one — or a few — sampled tokens. The resulting estimator can have large sampling error even when the teacher’s correction at that position is genuinely valuable. Feed a noisy estimate into the update and you can push the student in the wrong direction at a position you specifically selected because it mattered.

The authors formalize this in information geometry. At a fixed prefix, they derive a signal-to-noise decomposition of the one-sample reverse-KL gradient in Fisher geometry, and propose the information-efficiency ratio (IER) — the signal-to-noise ratio under the optimal scalar baseline that minimizes estimator variance. A useful theoretical aside in the appendix shows that IER⁻¹ is exactly the relative mean-squared error of the token-gradient estimate. Computing the exact IER over the full vocabulary is impractical, so the paper introduces a top-K candidate-set approximation, ranks tokens by approximate IER, and — crucially — combines that ranking with existing usefulness scores through soft logical OR and AND operators (IER-OR picks tokens that are useful or reliable; IER-AND picks tokens that are useful and reliable), while keeping the same sampled reverse-KL training objective.

The headline results

The experimental grid is admirably broad: verifiable math reasoning and open-ended medical reasoning, strong-to-weak distillation (post-trained teacher, same-size student) and big-to-small distillation, thinking-on and thinking-off modes.

For mathematics, the paper distills JustRL-Nemotron-1.5B into OpenMath-Nemotron-1.5B, and JustRL-Qwen3-4B into Qwen3-1.7B, training on the DAPO-Math-17k prompt pool for 50 rollout rounds and evaluating on AIME 2025/2026 and HMMT-Feb 2025/2026 with 32 samples per problem (Bayes@32 metric). For medical reasoning, ClinAlign-4B is distilled into Qwen3-4B over RaR-Medicine for 100 rollout rounds, scored on HealthBench and HealthBench Hard with gpt-oss-120B as grader — a model that itself carries a 0.6614 macro-F1 in HealthBench’s meta-evaluation.

Three findings stand out:

IER’s distribution is brutally heavy-tailed. Fewer than 0.1% of tokens have an IER score above 1 — meaning that for the overwhelming majority of positions, estimated noise exceeds signal. Full OPD, which supervises all tokens regardless of gradient reliability, is by this measure doing something close to optimal-by-accident at best, and adding noise at worst.

Extreme sparsity is viable. At a 0.1% token budget, IER as a standalone selector approaches full OPD on the Nemotron pair (AIME26: 58.9 vs. 59.9 for full OPD; HMMT26: 34.1 vs. 34.7) and exceeds full OPD on three of four benchmarks for the Qwen3 pair. Combined with usefulness scores at 1% budgets, sparse configurations match or beat full supervision — a 100×–1,000× reduction in the number of supervised token positions.

The two axes are genuinely complementary. The Jaccard-similarity analysis shows that IER’s top-10% selections often differ substantially from usefulness-based selectors even when their overall rank correlations are high — reliable tokens aren’t always the “important” ones, and vice versa. One striking detail: with only 10% of tokens, the TIP selector with IER even beats the teacher on HMMT26. And in the budget sweep from 0.1% to 80%, more supervision does not monotonically help; 1%–5% with usefulness+IER matches or outperforms full OPD in several settings.

The thinking-mode experiments carry a practical warning. With thinking-on (reasoning tokens included in selection), TIP alone underperforms full OPD and can even land below the un-distilled student — but TIP+IER-AND recovers the gains in most cases. Ignoring estimation noise is most dangerous exactly where the supervision is richest.

The honest caveats

Two limitations deserve emphasis, and the paper is refreshingly upfront about them. First, sparse supervision does not imply proportional compute savings here: the implementation still generates complete trajectories and then scores token-selection quality on top, costing 2.2%–2.5% extra step time and up to 2.17 GiB more peak GPU memory (0.92%) across eight GPUs. This is a study of where supervision pays off, not yet a production cost-cutter; the prize — skipping teacher scoring on ~99% of positions — goes to whoever combines IER-style selection with trajectory pruning or early termination. Second, gains vary across selectors and settings, and the practical IER is itself an approximation; the authors explicitly leave better token-importance understanding and training-objective design as open problems.

Why it matters

Distillation economics are becoming a first-order constraint across the industry — from open-weight labs training small models from frontier teachers to enterprises compressing models for on-device deployment. This paper gives the field a principled, composable answer to “which tokens deserve a teacher’s attention”: usefulness and reliability are orthogonal axes, and you need both. The metric is cheap to approximate, plugs into existing selectors rather than replacing them, and ships with code (Apache-style release on GitHub under BruceSheng1202/IER-OPD).

For medical AI in particular, the ClinAlign→Qwen3 results hint at something bigger: in domains where teacher scoring is expensive and errors are costly, being able to say “these 1% of tokens carry essentially all of the learnable signal” is a foundation for far cheaper, more targeted supervision pipelines. The paper’s own framing is more modest — a supplement to existing usefulness scores — but the direction is clear. The next generation of distillation recipes may spend as much effort deciding where not to look as where to look.