PANews
PANews|Sep 22, 2026 09:38
[Luofuli: MiMo-V2.6 May Have Completed One of the Largest Single Reinforcement Learning Trainings in the Open-Source Model Field] Luofuli, a member of Xiaomi's MiMo team, stated in a post that MiMo-V2.6, in terms of computational scale, might be one of the largest single reinforcement learning trainings conducted by an open-source model team to date. Despite limited computational resources, the team involved dozens of members and sustained efforts over an extended period to advance reinforcement learning expansion. The model's capabilities were formed during the mid-training phase and further unlocked through large-scale reinforcement learning. She mentioned that the team used MixRL training for medium-difficulty tasks that could be verified, such as coding, while separately training for tasks that were hard to verify, had ultra-long cycles, or were overly complex, and then merged capabilities through MOPD. To promote research in agent reinforcement learning, the team also released the Qwen model distilled from MiMo reinforcement learning trajectories, 7,000 diverse environments, and a complete reinforcement learning training framework.
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads