The Open-Weights Crown Changes Hands: Xiaomi's MiMo-V2.6-Pro Ties Grok 4.7 for Under $3 Million
Xiaomi's MIT-licensed MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index to become the strongest open-weight model in the world — after a $2.62 million reinforcement-learning run.
For years, the strongest open-weight AI models have come from specialist labs — DeepSeek, Z.ai, Moonshot AI. This week the crown moved to a company best known for phones and electric cars. On September 21, 2026, Xiaomi released the MiMo-V2.6 family, and its flagship, MiMo-V2.6-Pro, promptly scored 46 on the Artificial Analysis Intelligence Index: the highest rating of any open-weight model available, tying xAI’s freshly shipped Grok 4.7 and taking the title from Z.ai’s GLM-5.3 and Moonshot’s Kimi K3, both at 44. The jump within Xiaomi’s own lineup is even steeper — MiMo-V2.5-Pro managed just 26 on the same index.
The headline number that has the industry talking, though, is the bill. Xiaomi says it got there with a reinforcement-learning run costing roughly $2.62 million — under $3 million to close most of the gap to the closed frontier.
What shipped
The MiMo-V2.6 family has three tiers:
- MiMo-V2.6-Pro — a sparse mixture-of-experts model with 1.02 trillion total parameters, of which only about 42 billion activate per token.
- MiMo-V2.6-Flash — the lighter sibling at 309–310 billion total parameters and 15 billion active, trained through the same pipeline for around $850,000.
- MiMo-V2.6-Pro-UltraSpeed — a variant Xiaomi says delivers up to 20× output speed at identical quality, priced at ten times standard Pro.
All three handle text, image, speech, and video natively as input (text output), with a context window of one million tokens. The weights are published on Hugging Face under an MIT license — among the most permissive licenses in existence — alongside a 9-billion-parameter distill built on Qwen3.5, the full technical report, more than 7,000 reinforcement-learning environments, and the training code. One Bluesky reviewer called it “the most batteries-included tech report for a modern frontier model ever.”
The team behind it is led by Fuli Luo, a former DeepSeek researcher who now heads Xiaomi’s MiMo group, with “several dozen people” assigned to the RL effort, VentureBeat reported.
How it stacks up
On the overall Artificial Analysis ranking — which scores open and closed models side by side — MiMo-V2.6-Pro lands sixth, with every model above it proprietary:
| Model | AA Intelligence Index |
|---|---|
| Claude Fable 5.1 (Anthropic) | 53 |
| GPT-6 Astra (OpenAI) | 53 |
| Claude Opus 5 (Anthropic) | 51 |
| Muse Spark 1.3 (Meta) | 48 |
| GPT-5.6 Sol (OpenAI) | 47 |
| MiMo-V2.6-Pro (Xiaomi) | 46 |
| Grok 4.7 (xAI) | 46 |
| Gemini 3.8 Flash (Google) | 41 |
| DeepSeek V4.1 Pro (DeepSeek) | 36 |
That is a seven-point gap to the very top and a single point to the next proprietary model up. For an open-weights system you can download and self-host, that is a remarkably thin margin.
Xiaomi’s own benchmark card is more nuanced, showing clear strengths and one soft spot:
- Cybersecurity: 94.0 on CyberGym for Pro; Flash actually takes first place at 95.1, ahead of DeepSeek V4.1 Flash (88.1) and GLM-5.3 (84.5).
- Tool use: 76.9 on Toolathlon-verified, ahead of GPT-5.6 Sol (74.9) and just behind the Claude models.
- Economically relevant tasks: 1,673 on GDPVal — third overall, behind only Claude Fable 5.1 (1,735) and Claude Opus 5 (1,708).
- Software engineering: 71.9 on DeepSWE v1.1, against 74.0 for Claude Opus 5 and GPT-6 Astra.
- Visual coding: 72.3 on Xiaomi’s in-house test, above Claude Opus 5 (70.0), below GPT-6 Astra (82.2).
- The weak spot: long terminal sessions. Pro reaches 34.9 on Terminal Bench 4.0 versus 59.6 for GPT-6 Astra and 49.0 for Claude Opus 5. ExploitGym tells a similar story at 17.8, well behind the proprietary leaders. (On the older Terminal Bench 2.1, the model scores a much healthier 89.9.)
The economics: cents, not dollars
The cost picture is where MiMo-V2.6-Pro stands out most. Artificial Analysis puts its spend at roughly €0.11–0.13 per Intelligence Index task, placing it on the Pareto frontier of intelligence versus price. OfficeChai estimates that is roughly one-twentieth to one-sixtieth the cost of leading international models for the same work. Throughput is solid too: about 125 output tokens per second, 12th among 114 tested models.
The one flagged inefficiency is verbosity — the model burned 140 million output tokens to complete the index run, echoing the same token-consumption critique levelled at Grok 4.7 this week.
API pricing stays level with the previous generation: Pro at $0.435 per million input tokens and $0.87 per million output (with a 99% discount on cache hits); Flash at $0.14/$0.28. The models are available through Xiaomi’s AI Studio, MiMo Code, the MiMo Desktop app (which just left early access), Xiaomi’s own API platform, and OpenRouter.
“You Only RL Once”
Xiaomi credits the capability leap to scaled reinforcement learning on verifiable, complex tasks. In under six days, Flash and Pro each completed 30 RL steps across roughly 750,000 trajectories — $740,000 for Flash, $2.28 million for Pro. On the held-out DeepSWE v1.1 benchmark, Pro climbed from 58.4 to 72.6 points and Flash from 48.8 to 65.7 over the course of the run. Training used 1,568 samples per update, a fully asynchronous architecture, and context lengths up to one million tokens. Xiaomi says it streamed the production run live as it happened.
The distinctive methodological choice: rather than training each capability in isolation, the team ran a single unified pass spanning coding, general agents, visual, and cybersecurity tasks — an approach Xiaomi brands “You Only RL Once.” Releasing the environments and RL code alongside the weights means the claim is, in principle, reproducible.
From vibe coding to “Vibe World”
Xiaomi’s headline pitch is that coding ability now generalizes beyond software engineering — what the company calls “Vibe World.” The demonstrations range from generating 3D scenes with interaction logic, modeling objects in Blender, and controlling a Franka Panda robotic arm through multiple camera feeds, to frontend interfaces and slide decks from a one-line brief, video editing and scoring, and even composing an orchestral score for roughly ten instruments and converting it to MIDI unaided.
Two research case studies stand out. Working with Xiaomi’s materials scientists, the model designed a new metal-organic framework for binding PFAS “forever chemicals” — reviewing literature and patents, then computing binding strengths with open-source simulation tools. In the second, it helped formalize the main theorem of Li and Yorke’s classic “Period Three Implies Chaos” paper in Lean 4, producing more than 6,000 lines of proof code that the Lean kernel verified in full.
Why it matters
Three things make this release more than a leaderboard blip. First, the price of admission: if a hardware conglomerate can reach the top of the open-weights table for under $3 million of RL compute, the marginal cost of frontier-adjacent capability keeps collapsing — and the pricing pressure on US closed labs is now denominated in cents per task, not dollars. Second, the license: MIT weights plus training code and environments is as open as frontier AI gets, and it hands every researcher and enterprise a reproducible artifact rather than a marketing claim. Third, the source: Xiaomi is primarily a device maker. Its arrival at the top of the open-weights charts signals that frontier-adjacent model training is becoming a mainstream engineering discipline rather than the preserve of a handful of labs.
The caveats deserve their own line. The RL cost figure is vendor-reported, not an independent audit, and it may reflect internal compute subsidies unavailable to competitors on public cloud. The model still trails clearly on long-horizon terminal work and offensive security tasks, exactly the areas where agentic AI is heading. And the leaderboard is a snapshot: if Meta’s Muse Spark 1.3 or GPT-5.6 Sol drift higher at the next Artificial Analysis refresh while Pro holds at 46, Xiaomi’s claim becomes a moment rather than a trend.
But for now, the strongest open-weight model in the world comes from a phone maker — and you can download it, inspect it, and rebuild its training run yourself.