A 2B Model That Beats the 4B Class: OpenBMB Open-Sources MiniCPM5-2B With 131K Context and a Fully Open Data Stack
OpenBMB's MiniCPM5-2B averages 53.9 across 34 benchmarks, outscoring every 4B-class rival in its comparison set, and ships with 131K context, hybrid thinking, and the entire UltraData training stack under Apache 2.0.
The small-model race just produced its most aggressive opening move of the autumn. On September 7, 2026, OpenBMB — the open lab spun out of Tsinghua University’s Big Model Base research group — released MiniCPM5-2B, a dense 2.5-billion-parameter language model that not only claims state-of-the-art status within the 2B class but outright outperforms every larger model in its published comparison set, including Alibaba’s Qwen3.5-4B. And in what has become the lab’s signature move, it shipped the entire training stack — pre-training corpora, SFT data, RL data, and intermediate checkpoints — alongside the weights, all under the permissive Apache 2.0 license.
What the numbers say
MiniCPM5-2B is the second model in the MiniCPM5 series, following the MiniCPM5-1B released earlier this year. On paper it is deliberately unexotic: a standard LlamaForCausalLM architecture with 42 layers, grouped-query attention using 16 query heads and 2 KV heads, 2,516,756,480 total parameters (1.98B non-embedding), and a native context length of 131,072 tokens — an unusually large window for a model this size.
Across the 34 benchmarks OpenBMB reports, the model averages 53.9. That figure beats its direct 2B-class peers decisively — LiquidAI’s LFM2.5-2.6B averages 33.2, Qwen3.5-2B scores 28.0, and Google’s Gemma-4-E2B-it manages 24.6 — but the more striking result is the cross-class comparison: the best-scoring model in the entire 4B-class reference group, Qwen3.5-4B, averages just 51.1. IBM’s granite-4.2-3B lands at 42.7, Nemotron-3-Nano-4B at 32.6, and Gemma-4-E4B-it at 31.2. A 2.5B model outscoring a 4B model on an aggregated 34-benchmark suite is exactly the kind of result that reshapes buyers’ intuitions about what a “small” deployment can do.
Independent scoring broadly corroborates the claim. Artificial Analysis, whose methodology differs from OpenBMB’s internal reproductions, scored MiniCPM5-2B at 15 on its Intelligence Index v4.2 — the highest of any open-weights model under 4B total parameters. The directional finding — that this model tops its weight class — survives even when measured by a party that didn’t train it.
Where the margins come from
The sub-4B leaderboard is one thing; the individual benchmark rows are where the release gets interesting.
Mathematics is the loudest column. MiniCPM5-2B scores 86.5 on both AIME 2025 and AIME 2026, versus 78.8/82.7 for Qwen3.5-4B and 79.4/83.5 for granite-4.2-3B. It clears 94.6 on MATH-500. Competition-math performance at that level from a 2B dense model would have been a frontier result for models several times larger only two years ago.
Coding tells a similar story. The model posts 69.1 on LiveCodeBench v6 — nearly 13 points above Qwen3.5-4B’s 56.4 — and 68.0 on the LCB-Pro easy split. On the harder agentic-coding terrain it still leads its size class: 46.4 on SWE-bench Verified, where the next-best 2B-class model manages 6.0 and Qwen3.5-4B itself scores only 33.6. On SWE-bench Pro the larger Qwen model regains the lead (28.2 vs 14.4), a reminder that complex multi-file repo work still rewards raw parameter count.
Tool use and long context are the two areas OpenBMB is clearly positioning for the on-device agent era. The model scores 97.1 on τ²-Bench Telecom, a customer-service tool-use benchmark, ahead of every model in the table. On NoLiMa, a long-context benchmark that tests associative recall under lexical masking, it scores 68.1 while most competitors collapse to single digits — Qwen3.5-4B manages 43.5, LFM2.5-2.6B just 0.7. The native 131K context window is clearly not decorative.
Not everything is blue on the scoreboard, and it’s worth being precise about that. Qwen3.5-4B retains the lead on MMLU-Pro (78.0 vs 70.8), MATH-500 (99.0 vs 94.6), GPQA-Diamond (77.1 vs 70.2), IFBench, SWE-bench Pro, and Terminal-Bench v2.1. The honest summary is that MiniCPM5-2B wins the aggregate through decisive leads in code reasoning, math reasoning, tool use, long-context, and agentic tasks, while ceding some breadth in knowledge-intensive and multi-step terminal work.
The training recipe: RL teachers distilled on-policy
Under the hood, MiniCPM5-2B is the flagship demonstration of OpenBMB’s “UltraData Tiered Data Management” methodology, documented in the lab’s technical reports. Training runs through three stages: base training, mid-training, and post-training, with the post-training phase itself split into SFT, RL, and a stage the lab calls OPD — On-Policy Distillation.
The sequence is unusual and worth understanding. First, 400B tokens of deep-thinking SFT data establish reasoning and general chat behavior. Then OpenBMB trains not one but 16 specialized RL teacher models — covering math, code, agentic tasks, and writing, five of them agentic experts — using a critic-based algorithm from the lab’s JustRL II work that substantially improves training stability. Finally, OPD distills those teachers back into the single 2B release model: at each response position, a full-vocabulary reverse KL divergence between student and teacher logits serves as the advantage estimate, replacing verification-based advantage, and the distillation prompts simply reuse the RL teachers’ own training data.
The payoff is quantified: RL + OPD improves reasoning and general capabilities by an average of 10.96 points and agentic capabilities by 6.96 points over the SFT baseline. In other words, a meaningful fraction of the benchmark lead comes not from the pre-training run but from this multi-teacher distillation pipeline — a recipe other small-model teams will now have to treat as the reference implementation, especially since the RL data itself (80K+ samples) is open-sourced as UltraData-RL-2609.
The openness is arguably the release’s defining feature. Alongside the final BF16 model, OpenBMB published the SFT-only checkpoint, the mid-training checkpoint, the base checkpoint, GGUF builds for llama.cpp/Ollama/LM Studio, an MLX 4-bit build for Apple Silicon, a GPTQ quantization, and a DSpark speculative-decoding draft model for accelerating inference without changing outputs. The data side includes UltraX (web pre-training data), UltraData-Code with L0–L3 tiered code data management, UltraData-SFT-Agent with 500K agent training samples, and the RL corpus. It is close to a complete open recipe for reproducing a competitive small model.
Built for the edge, deployed everywhere
Deployment ergonomics follow from the architecture choice. Because MiniCPM5-2B is a standard LlamaForCausalLM, mainstream inference engines load it directly — no custom kernels, no forked model code. OpenBMB ships documented paths for vLLM (vllm>=0.21), SGLang (recommended for tool calling, with a native minicpm5 parser that converts the model’s XML-style tool calls to OpenAI-compatible tool_calls), Transformers, llama.cpp, Ollama, LM Studio, and MLX. Fine-tuning cookbooks cover TRL+PEFT, LLaMA-Factory, ms-swift, and unsloth.
There is also a distinctly Chinese-ecosystem dimension to the release: through the FlagOS open-source system stack, the model has already been adapted to nine different AI chip platforms — Nvidia, Hygon, Metax, Iluvatar, Zhenwu, MetaX, Mthreads, Kunlunxin, Huawei Ascend, and ARM v9 — and published on the FlagRelease platform for cross-chip deployment. For an industry watching compute fragmentation closely, a strong small model that runs “develop once, deploy across chips” is a quietly strategic artifact.
Why it matters
The sub-4B segment is where the volume is. It is the tier that runs on phones, laptops, robots, vehicles, and edge boxes — the devices that cannot afford a frontier model’s latency or power budget. For most of the past year that segment has been dominated by Qwen’s dense small models, which combined strong per-parameter efficiency with Alibaba’s distribution muscle. A challenger that outperforms Qwen3.5-4B on aggregate with roughly 60% of the parameters — while also publishing the full data and RL pipeline — applies real competitive pressure precisely where unit economics are decided.
The timing also lands amid a broader industry conversation about “model fatigue,” with IT buyers telling outlets like CNBC that the relentless cadence of frontier releases has become exhausting. Small-model releases of this kind are the counter-argument: rather than asking enterprises to re-qualify a new flagship every few weeks, they push capability down into footprints that already exist on devices people own. A 2B model with credible AIME scores, SWE-bench Verified in the mid-40s, 131K context, and a 97 on tool-use benchmarks is not a toy — it is a plausible daily driver for local assistants and coding agents.
Caveats remain. OpenBMB’s internal reproductions cover most scores (with Artificial Analysis supplying the dagger-marked ones), and self-reported comparison sets always invite scrutiny of what was left out. Hybrid-thinking small models also tend to trade latency for accuracy when the thinking mode engages, and the model’s own disclaimer is blunt that outputs may be inaccurate or biased and that sensitive-domain answers are not expert-reviewed. Qwen’s next small-model refresh will presumably answer the challenge directly.
But as a single data point about where the efficiency frontier sits in September 2026, MiniCPM5-2B is hard to ignore: more capability per parameter than anything else in the open sub-4B field, and a release posture — full data, full checkpoints, full recipe — that turns a model launch into a methodology transfer.