律动BlockBeats|Sep 22, 2026 04:12
[Luofuli: The development challenges of MiMo-V2.6 have surpassed those of DeepSeek R1 that I participated in]
Beating AI Newsflash: Xiaomi MiMo lead Luofuli further elaborated on this round of reinforcement learning following the release of MiMo-V2.6. She stated that she had previously participated in DeepSeek R1, but in her view, the research innovations and engineering challenges behind MiMo-V2.6 have already exceeded those of R1. She also explained why MiMo simultaneously utilizes MixRL and MOPD. MixRL integrates verifiable tasks such as code, general agents, vision, and network security into the same round of reinforcement learning for joint training. MOPD, on the other hand, handles ultra-long, hard-to-verify, or reward-subjective tasks by training them separately first and then integrating these capabilities back into the main model. Tasks like games and 3D have excessively long runtime and are difficult to automatically determine correctness. If combined with tasks like code in the same round of RL, it would significantly slow down training. Therefore, MiMo trains such tasks separately and then reintegrates the learned capabilities into the main model via MOPD. [Original link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink