SAVRN
Search Contact SAVRN

Organization

Xiaomi MiMo

XiaomiMiMo

Models in Library2
Datasets in Library0
Models on Hugging Face28
Followers4.2k

Models

Model · Text generation

MiMo-V2.6-Flash-RL

Xiaomi MiMo

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers

Model · Text generation

MiMo-V2.6-Pro-RL

Xiaomi MiMo

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn…

Open weights mit 524.1B parameters 1,048,576 tokens transformers