Abstract
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
Community
Technical report of MiMo V2.6 series.
Blog: https://mimo.xiaomi.com/mimo-v2-6
HF weights: https://huggingface.co/collections/XiaomiMiMo/mimo-v26
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Intern-S2-Preview: Scientific Agentic Foundation Model (2026)
- Learning Generalizable Behaviors for Terminal Agents (2026)
- Pistis Technical Report (2026)
- QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents (2026)
- ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning (2026)
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (2026)
- DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper