← All posts / Models

You Only RL Once: Xiaomi Open-Sources MiMo-V2.6 Pro and Flash at Claude-Opus-Level Agentic Scores

Xiaomi has released the MiMo-V2.6 series under MIT license: a 1.02T-parameter omnimodal flagship scoring 71.9 on DeepSWE and 31.6 on Agents' Last Exam, a 309B Flash sibling, a distilled 9B checkpoint, and the RL training environment itself.

You Only RL Once: Xiaomi Open-Sources MiMo-V2.6 Pro and Flash at Claude-Opus-Level Agentic Scores

On September 21 at roughly 15:39 UTC, the XiaomiMiMo organization on Hugging Face quietly flipped three repositories public: MiMo-V2.6-Pro-RL, MiMo-V2.6-Flash-RL, and MiMo-V2.6-Distill-Qwen-9B. No keynote, no livestream, no countdown. The weights of the model the company had been live-streaming through its reinforcement-learning run for the past week — entropy curves, benchmark checkpoints, even a disclosed VRAM crash — were suddenly just there, downloadable under an MIT license.

It closes one of the more unusual release cycles in recent open-weights history. Most frontier labs publish a curated technical report after shipping. Xiaomi spent the week before release letting anyone watch the model get better in public. Now the artifacts are in the open: two sparse mixture-of-experts models, a distilled student, and enough architectural detail to serve them at home.

What shipped

The flagship, MiMo-V2.6-Pro-RL, is a sparse MoE with 1.02 trillion total parameters and 42 billion activated — the same scale as its V2.5 predecessor, but now natively omnimodal. Text, image, video, and audio run through a single model with a 1M-token context window, fed by a 681M-parameter MiMo ViT vision encoder (28 layers, 24 sliding-window plus 4 full-attention) and a 308M AudioTokenizer paired with a 127M audio patch encoder that packs four frames per patch. A 5-layer multi-token-prediction drafter predicts 7 subsequent tokens per forward pass for speculative decoding.

The sibling, MiMo-V2.6-Flash-RL, keeps the omnimodal design but compresses it: 309B total parameters, 15B activated, 256 routed experts of which 8 fire per token, and the same 1M-token context. Both models share the hybrid backbone — 70 layers with 60 sliding-window-attention blocks and 10 global-attention blocks in the Pro variant — with no shared experts in the sparse FFNs.

The third release is easy to overlook: MiMo-V2.6-Distill-Qwen-9B, a 9B student distilled from the flagship, plus the RL training environment itself. For a community that has watched labs ship weights but keep the scaffolding closed, publishing the environment is arguably the more radical half of the drop.

The scores

The evaluation table in the model card is blunt about the ambition. On agentic coding, MiMo-V2.6 Pro scores 71.9 on DeepSWE v1.1 — within 2.1 points of Claude Opus 5 (74.0), ahead of GPT-5.6 Sol (73.0) and Claude Fable 5 (70.0). The jump from its own predecessor is the real story: MiMo-V2.5 Pro managed just 19.0 on the same benchmark. One generation, 52.9 points.

General agents tell the same story. On Toolathlon-Verified the Pro model hits 76.9 against Opus 5’s 80.6. On AutomationBench v1.0.6 it scores 53.1 — ahead of every closed competitor listed, including Opus 5 at 50.3. On Agents’ Last Exam it ties Opus 5 at exactly 31.6. Terminal Bench 2.1: 89.9, a hair over Opus 5’s 89.1. The gap narrows or flips on some suites — ExploitBench puts Opus-class models at 78.5 while the Pro lands at 47.9, and Terminal Bench 4.0 remains a closed-model stronghold at 49.0 versus 34.9 — but the pattern is unmistakable: on agentic workloads, the open frontier has closed to within rounding distance.

Third-party measurement backs the claim. Artificial Analysis scored the Pro model 46 on its Intelligence Index, the highest ever recorded for an open-weight model — territory previously held by the V2.5 Pro and Moonshot’s Kimi K2.6 at 54 on the older scale. On Xiaomi’s internal harness, DeepSWE lands at 72.57 (Pro) and 65.68 (Flash).

You Only RL Once

The interesting engineering is in how the capability was built. The team’s framing is “scaling reinforcement learning toward self-improvement,” and the model card describes a genuinely unusual recipe:

  • One mixed RL run across coding, general agents, visual, and cybersecurity domains — “You Only RL Once” — rather than separate per-domain post-training passes. Tasks from multiple harnesses share each batch so strategies transfer to harnesses the model never saw in training.
  • Fully asynchronous GRPO at extreme batch scale: 1,568 prompts × 16 rollouts per step, billions of tokens per update.
  • Groupwise Agentic Grading as a self-improvement loop. Binary pass/fail can’t rank solutions that all pass, so Xiaomi scaled the grader too: Groupwise Reward Synthesis builds task-specific rubrics offline from contrasting rollouts, and Groupwise Advantage Redistribution ranks passing trajectories online, shifting advantage toward higher-quality solutions. Judged against the policy’s own samples, the loop steers the model toward shorter paths and fewer tokens per task.
  • Aligned RL from cold start — the model begins from self-correction, rewriting its own misaligned turns into grounded next steps — with environment hardening, adversarial screening, and verifier cross-checks standing guard against reward hacking.
  • MOPD2, a multi-prefix multi-teacher on-policy distillation stage that runs after mixed RL, combining autonomous student rollouts with prefix-conditioned teacher and SFT rollouts so hard-to-verify tasks still get trained decision points without regenerating entire histories.

The cybersecurity numbers suggest the mixed run paid off in places pure code RL doesn’t usually reach: CyberGym at 94.0/95.1 (Pro/Flash) and MiMo Cyber Bench at 80.2/77.2, against V2.5 Pro’s 40.0 and 0.0.

Running it

Deployment recipes are published for both SGLang and vLLM. The SGLang reference config for the Pro model is not for the faint of heart — tensor parallel 16, expert parallel 16, DeepEP all-to-all, two nodes, EAGLE-style multi-layer speculative decoding — but the Flash variant’s 15B active parameters put it within reach of serious single-node setups, and the pre-built vllm/vllm-openai:mimov25-cu129 image carries over. Recommended sampling: temperature 1.0, top-p 0.95. The models are also served through Xiaomi’s own Open Platform API, AI Studio, the MiMo Code terminal agent, MiMo Desktop, and OpenRouter. An UltraSpeed variant on Xiaomi’s platform claims up to 20× inference throughput.

Why it matters

Two things make this release more than another weights dump. First, the MIT license with no usage restrictions lands a frontier-adjacent agentic model into commerce-ready territory just as enterprise AI startups — Harvey, Abridge, Decagon, Ramp — are publicly pivoting to open weights because closed-model token economics crushed their margins. A 46-index omnimodal agent that ties Claude Opus 5 on Agents’ Last Exam is now a build-on-it option, not a research artifact.

Second, the release completes an experiment in transparency. The week of live RL streaming was already a first; shipping the RL environment alongside the weights means the community can attempt the loop, not just the model. The predecessor’s track record helps: MiMo-V2.5 Pro built genuine adoption in local-LLM circles, and Reddit’s reaction to V2.6 has been the expected mix of benchmark excitement and “real-world performance may vary” caution.

The open-weights race has been a game of leapfrog between Chinese labs — Kimi, Qwen, DeepSeek, StepFun, now Xiaomi again. As of September 21, the frontrunner badge belongs to MiMo-V2.6 Pro. How long it holds is anyone’s guess; the pace suggests weeks, not months.