Interestana
Home/News/OpenBMB Releases MiniCPM5-2B Dense Model for On-Device Use
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

OpenBMB Releases MiniCPM5-2B Dense Model for On-Device Use

OpenBMB has released MiniCPM5-2B, the second checkpoint in its MiniCPM5 series and a successor to MiniCPM5-1B. This model is a dense causal language model featuring 2,516,756,480 parameters, with 1,981,982,720 parameters located outside the embedding layer. It incorporates 42 layers and utilizes grouped-query attention with 16 query heads and 2 key/value heads. MiniCPM5-2B boasts a native context window of 131,072 tokens. Its architecture is based on the standard LlamaForCausalLM, allowing for seamless loading by mainstream engines without requiring custom kernels or model-code forks. The model's weights are licensed under Apache 2.0 and are compatible with various deployment platforms including vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, MLX, and FlagOS, indicating its suitability for on-device applications.

In benchmark evaluations, MiniCPM5-2B was compared against models of a similar size class, including LFM2.5-2.6B, Qwen3.5-2B, and Gemma-4-E2B-it. For broader context, it was also benchmarked against larger models such as Qwen3.5-4B, granite-4.2-3B, Nemotron-3-Nano-4B, Gemma-4-E4B-it, and LFM2.5-8B-A1B. Across 34 benchmark tests, MiniCPM5-2B achieved an average score of 53.9. This performance surpasses the best baseline in its size class, Qwen3.5-4B, which scored 51.1, followed by granite-4.2-3B at 42.7 and LFM2.5-2.6B at 33.2. The model demonstrated particular strength in code reasoning, scoring 69.1 on LiveCodeBench v6, a significant improvement over the baseline's 56.4, and achieving 46.4 on SWE-bench Verified, compared to 33.6. Tool use also showed substantial gains, with scores of 97.1 on τ²-Bench Telecom and 66.6 on BFCL v4, though it scored 20.8 on τ³-Bench Banking against a baseline of 6.8.

Regarding long context capabilities, MiniCPM5-2B showed mixed results. It achieved a strong score of 68.1 on NoLiMa, outperforming the baseline's 43.5. However, on AA-LCR, it scored 59.0, slightly below the baseline of 61.0, and on LongBench v2, it scored 43.7, compared to 47.3. In general knowledge tasks, where larger models typically excel, MiniCPM5-2B scored 70.8 on MMLU-Pro, falling short of the baseline's 78.0, and achieved 8.9 on Humanity’s Last Exam, versus 9.9. OpenBMB distinguishes benchmark rows sourced from Artificial Analysis from those reproduced internally. The training methodology for MiniCPM5-2B involved Supervised Fine-Tuning (SFT), followed by Reinforcement Learning (RL), and then on-policy distillation. The training process adheres to the UltraData tiered data management method, as detailed in original research. Base training encompasses stable and decay phases, with mid-training adjustments to align the model with the target data distribution. Post-training begins with 400 billion tokens of deep-thinking SFT, subsequently training specialized RL teachers for tasks such as mathematics, coding, and agentic operations.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next