← All posts / Models

Training in Public: Xiaomi Streams MiMo-V2.6's Live RL Run, Logs and Failures Included

Xiaomi is streaming the raw reinforcement-learning metrics of MiMo-V2.6 Pro and Flash straight from the trainer's logs — entropy, pass rates, infra errors, even a VRAM crash notice — a level of openness no frontier lab has tried.

Training in Public: Xiaomi Streams MiMo-V2.6's Live RL Run, Logs and Failures Included

Frontier AI labs publish polished technical reports after a model ships. Xiaomi just did the opposite: with the MiMo-V2.6 series still mid-training, the company has opened a public dashboard that streams the raw reinforcement-learning metrics of two runs — mimo-v2.6-pro and mimo-v2.6-flash — directly from the trainer’s logs, updating live as the runs progress.

The page, live at mimo.xiaomi.com/rl/ since roughly 04:00 UTC on September 16, shows policy entropy, policy-gradient loss, gradient norms, mean trajectory rewards, and pass-rate histograms re-plotting in real time. The runs are agentic RL: prompts are sorted into five categories — code, general, cyber, visual, and chat — and the models practice in sandboxed environments, with a live counter of environments in flight. Benchmark checkpoints arrive as the run progresses, with the team posting offline evaluation results to a notices feed as they land.

What the stream actually shows

The dashboard’s pinned headline metric is dynsam/avg@n — for each prompt sampled in a step, the fraction of its n attempts that succeed, averaged over prompts. Around it sit the internals that normally never leave a lab: per-token policy entropy (actor/entropy_loss), the clipped policy-gradient objective, global gradient norm before clipping, KL divergence between the inference engine and the trainer on the same tokens, average agent turns per trajectory, and the share of sequences lost to infrastructure failures.

The benchmark feed already tells a story. On DeepSWE v1.1 (mini-swe-agent, avg@3), the Flash run has climbed from 48.67 at step 1 to 60.77 by step 12, with the Pro run moving from 58.41 at step 1 to 63.72 at step 10. These are checkpoint evaluations of a model whose weights do not exist publicly yet — the community is literally watching a model get better at agentic coding in real time.

The notices feed is where the transparency gets unusual. One entry, timestamped 20:08 UTC on September 16, reads plainly: “the mimo-v2.6-pro run is restarting due to a vram issue on one node.” No PR spin, no silent retry — a hardware failure in a frontier-scale training run, disclosed in public minutes after it happened. Another notice confirms the team pushes fresh DeepSWE results as offline evaluations complete.

Why this matters

The industry norm is a closed loop: train in secret, evaluate in private, then publish a curated technical report alongside the release. Even the most open labs — those that ship weights and detailed papers — publish only after the fact, when the training narrative can be shaped in hindsight. Xiaomi is inverting that: the “report” is being written live, in trainer logs, before anyone knows how the story ends.

The gesture lands in a specific context. MiMo-V2.5 Pro, the current generation, built a genuine following in the open-weights community as a 1T-parameter MoE that competes with far more expensive APIs on agentic coding and long-context work. Expectations for V2.6 are high precisely because V2.5 over-delivered. By streaming the run, Xiaomi converts that goodwill into live engagement — and makes a verifiable claim that’s hard to fake: either the numbers trend up in public, or they don’t.

There’s also a research angle. RL training curves are among the least-documented aspects of frontier model development. Entropy collapse, reward-hacking drift, and infrastructure error rates during large agentic RL runs are discussed mostly in anecdotes. A live public stream of these exact quantities — including the ugly ones, like the share of sequences lost to infra failures and average rollout staleness — gives the research community a rare longitudinal view of a real big run.

The community reaction

The dashboard hit r/LocalLLaMA within hours and drew exactly the discussion you’d expect: excitement about the transparency, tempered by informed skepticism about what a training curve can and cannot prove. Commenters noted that post-training is more than RL — data curation and mixture decisions happen offline and stay invisible — and that a rising pass-rate curve doesn’t guarantee the final checkpoint clears the benchmarks that matter. The recurring theme: fascinating to watch, but weights or it didn’t happen.

That’s the right frame. The stream is a commitment device, not a product. The about text says only: “We are streaming our RL big runs. The mimo-v2.6 series is coming soon.” No release date, no benchmark claims, no pricing. Just a footer note — “Open is what we value.” — and two live curves.

The catch

It’s worth stating plainly what this is not. The dashboard shows what Xiaomi chooses to show: training-side metrics and self-reported benchmark checkpoints, with no external evaluation yet. A live curve can’t rule out cherry-picked data mixtures, and agentic-RL pass rates on curated prompt sources don’t always transfer to messy real-world tasks. The DeepSWE numbers arrive on the team’s own schedule, posted “as the offline evaluation results come out.”

And one run already restarted mid-stream — which is either an honest disclosure or a reminder that the stream shows a smoothed version of a messy physical process. Probably both.

Outlook

For anyone tracking the open-weights frontier, the watch item is simple: whether MiMo-V2.6 lands with actual open weights, and whether the final published benchmarks trace cleanly back to the curves now streaming in public. If they do, Xiaomi will have demonstrated something new — that radical process transparency can coexist with frontier-adjacent results. If they don’t, the dashboard becomes a case study in transparency theater.

Either way, for the next however many days the run lasts, you can watch a frontier-lab-grade RL run think — entropy rising, pass rates ticking up, sandboxes crashing — at a URL anyone can visit. That alone breaks the genre’s rules.