SAVRN
Search Contact SAVRN

Open-weight model · Text generation

MiMo-V2.6-Pro-RL

by Xiaomi MiMo XiaomiMiMo/MiMo-V2.6-Pro-RL

MiMo-V2.6-Pro-RL is an open-weight model for text generation from Xiaomi MiMo, released under MIT License. It has 524.1B parameters and a 1,048,576-token context. At 16-bit it needs about 1257.9 GB of GPU memory, which fits on 5x MI325X from $10.00 an hour; at 4-bit, 314.5 GB on 2x MI300X from $3.70, at the lowest prices in the SAVRN Index.

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series.

Parameters524.1B
Context1,048,576
Weights573.5 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve MiMo-V2.6-Pro-RL (524.1B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1048.2 GB 1257.9 GB 5x MI325X (256 GB)
Vultr
$10.00 5x MI355X $12.95 · 7x MI300X $12.95
8-bit 524.1 GB 628.9 GB 3x MI325X (256 GB)
Vultr
$6.00 4x MI300X $7.40 · 3x MI355X $7.77
4-bit 262.1 GB 314.5 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

MiMo-V2.6-Pro-RL on every accelerator the SAVRN Index prices, at every precision

Model Card

By Xiaomi MiMo, published under mit, revision 438e9ce7e34b.

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn…

Read Xiaomi MiMo's full model card





Community
WeChat Group  |  Discord  |  Telegram  |  Reddit


MiMo-V2.6-Pro-RL

Scaling Reinforcement Learning Toward Self-Improvement

Technical Report

1. Introduction

MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include:

  • Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs.
  • You Only RL Once: One mixed RL run across coding, general agents, visual, and cybersecurity — not separate per-domain runs. Tasks and multiple harnesses are mixed in the same batch so capabilities reinforce each other and strategies transfer to harnesses never seen in training.
  • Scaling RL Compute: Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches — 1,568 prompts × 16 rollouts per step, billions of tokens per update.
  • Groupwise Agentic Grading (Self-Improvement Loop): Binary pass/fail cannot rank passing solutions, so the reward signal itself is scaled. An agentic grader compares rollouts within each group: Groupwise Reward Synthesis (GRS) builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes; Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.
  • Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
  • Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix), reusing histories from teacher trajectories and SFT demonstrations so decision points train without regenerating preceding turns — extending capabilities to hard-to-verify tasks.

Model Summary

  • Architecture: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters
  • Context Length: 1M tokens
  • Modalities: Text, Image, Video, Audio
  • Vision Encoder: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
  • Audio Encoder: 308M AudioTokenizer + 127M audio patch encoder
  • Multi-Token Prediction (MTP): 5-layer speculative decoder

Figure 1. MiMo-V2.6 architecture.

2. Downloads

Model Download
MiMo-V2.6-Pro-RL HuggingFace · ModelScope
MiMo-V2.6-Flash-RL HuggingFace · ModelScope

3. Evaluation Results

Benchmark MiMo-V2.6 Pro MiMo-V2.6 Flash MiMo-V2.5 Pro Claude Opus 5 GPT-5.6 Sol Claude Fable 5
Code Agent
DeepSWE v1.1 71.9 67.9 19.0 74.0 73.0 70.0
ProgramBench 26.5 26.0 12.5 37.0 25.0 33.0
MiMo Code Bench 63.2 61.2 40.4 68.6 59.3 -
General Agent
AutomationBench v1.0.6 53.1 52.3 16.0 50.3 45.8 46.2
Toolathlon-Verified 76.9 73.6 49.1 80.6 74.9 77.9
GDPval-AA 2.1 1673 - 1107 1708 1588 1595
Agents’ Last Exam 31.6 27.6 13.2 31.6 30.8 25.7
Terminal Bench 4.0 34.9 28.8 1.5 49.0 39.9 42.4
Terminal Bench 2.1 89.9 87.6 65.2 89.1 88.8 84.3
OSWorld-Verified 82.0 80.8 - 83.4 83.0 86.0
JobBench 62.0 61.2 25.0 65.7 45.4 57.4
Cybersecurity
CyberGym 94.0 95.1 40.0 - - -
MiMo Cyber Bench 80.2 77.2 0.0 - - -
ExploitGym 17.8 6.0 0.2 22.1 30.3 28.4
ExploitBench 47.9 25.3 16.6 70.0 78.5 78.0
SEC Bench Pro 66.3 47.5 17.7 - 79.1 -
Visual Agent
MiMo VisualCoding 72.3 71.5 - 70.0 73.4 69.1

4. Model Architecture

LLM Backbone

Component MiMo-V2.6-Pro-RL
Layers (Total / SWA / GA) 70 / 60 / 10
Hidden Size 6144
SWA Heads (Q/KV) 128 / 8
GA Heads (Q/KV) 128 / 8
Head Dimensions (QK / V) 192 / 128
Sliding Window Size 128
Routed Experts (Total / Activated) 384 / 8
Max Context Length 1M
MTP / Speculative Decoder 5 SWA layers, window 1024

The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.

Vision Encoder (MiMo ViT)

Configuration Value
Layers (Total / SWA / GA) 28 / 24 / 4
Hidden Size 1280
Attention Heads (Q / KV) 32 / 8
Head Dimension 64
Patch Size (T × H × W) 2 × 16 × 16
Sliding Window (Left / Right) 64 / 64
Spatial Merge Size 2 × 2
Parameters 681M

Audio Encoders

AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).

Speculative Decoder

5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.

5. Deployment

For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.

SGLang

sglang serve \
  --trust-remote-code \
  --model-path XiaomiMiMo/MiMo-V2.6-Pro-RL \
  --tp 16 \
  --dp 2 \
  --enable-dp-attention \
  --mm-enable-dp-encoder \
  --ep 16 \
  --moe-a2a-backend deepep \
  --moe-dense-tp-size 1 \
  --mem-fraction-static 0.7 \
  --max-running-requests 128 \
  --chunked-prefill-size 32768 \
  --page-size 64 \
  --swa-full-tokens-ratio 0.3 \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --enable-multi-layer-eagle \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --host 0.0.0.0 \
  --port 30000 \
  --nnodes 2 \
  --node-rank <node-rank> \
  --dist-init-addr <node0-ip>:20000

vLLM

Follow the vLLM MiMo-V2.5 recipe. Pre-built image: docker pull vllm/vllm-openai:mimov25-cu129.

vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL \
  --tensor-parallel-size 8 \
  --trust-remote-code \
  --gpu-memory-utilization 0.95 \
  --max-model-len auto \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --enable-auto-tool-choice \
  --generation-config vllm

Recommended sampling: temperature=1.0, top_p=0.95.

Also available in AI Studio, MiMo Code, Xiaomi MiMo Desktop, Xiaomi MiMo Open Platform API, and OpenRouter.

Citation

@misc{mimo2026v26pro,
  title={MiMo-V2.6-Pro-RL},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}

Contact

For questions or feedback, reach us at [email protected] or join our community:

Configuration

Architecture
MiMoV2ForCausalLM
Context length (tokens)
1,048,576
Layers
70
Hidden size
6,144
Feed-forward size
16,384
Attention heads
128
Key/value heads
8
Head dimension
192
Vocabulary size
152,576
Routed experts
384
Experts active per token
8
Sliding window (tokens)
128
RoPE base
10,000,000
Model type
mimo_v2
Quantization
fp8

Identity and Version

Repository
XiaomiMiMo/MiMo-V2.6-Pro-RL
Publisher
Xiaomi MiMo
Task
Text generation
Modality
Text
Library
transformers
Parameters
524.1B parameters
Languages
en, zh
Revision
438e9ce7e34bc0188b878e6d6c76419b40d5427d
First published
2026-09-21
Last updated
2026-09-22

Files and Weights

155 files, 573.5 GB in total. The weights are 133 files totalling 573.5 GB in pt, safetensors.

Weights133 files · 573.5 GB
Configuration11 files · 15.3 MB
Tokenizer5 files · 15.9 MB
Documentation1 file · 9.8 KB
Other4 files · 3.5 MB
Repository1 file · 1.8 KB
Every file
FileTypeSizeSHA-256
audio_tokenizer/model.safetensorsWeights1.9 GB 077033345d80
dflash/dflash_draft_model.safetensorsWeights5.5 GB 39208e3dd452
dflash/mask_embedding.ptWeights14.0 KB 436ad0885387
model_mtp.safetensorsWeights2.5 GB eba334c8613d
model_pp0_ep0_shard0.safetensorsWeights34.4 GB 457fbf20a75a
model_pp0_ep0_shard1.safetensorsWeights2.0 GB 4f5be09d843b
model_pp0_ep100_shard0.safetensorsWeights4.2 GB 47403b29d884
model_pp0_ep101_shard0.safetensorsWeights4.2 GB 45fb8bbbc1f9
model_pp0_ep102_shard0.safetensorsWeights4.2 GB 7db63a2f918a
model_pp0_ep103_shard0.safetensorsWeights4.2 GB 812c35bd8eda
model_pp0_ep104_shard0.safetensorsWeights4.2 GB afb2cdf18078
model_pp0_ep105_shard0.safetensorsWeights4.2 GB 7cfe0e99ba9d
model_pp0_ep106_shard0.safetensorsWeights4.2 GB 2476b54862ae
model_pp0_ep107_shard0.safetensorsWeights4.2 GB 998e0c383b3e
model_pp0_ep108_shard0.safetensorsWeights4.2 GB 5d24c1b87af2
model_pp0_ep109_shard0.safetensorsWeights4.2 GB 7a46b4a647d9
model_pp0_ep10_shard0.safetensorsWeights4.2 GB 24d80d06ac43
model_pp0_ep110_shard0.safetensorsWeights4.2 GB 22bde5b3769c
model_pp0_ep111_shard0.safetensorsWeights4.2 GB e777a871c219
model_pp0_ep112_shard0.safetensorsWeights4.2 GB 86ec11f215f9
model_pp0_ep113_shard0.safetensorsWeights4.2 GB 79d80c98115d
model_pp0_ep114_shard0.safetensorsWeights4.2 GB 0ae9e102623b
model_pp0_ep115_shard0.safetensorsWeights4.2 GB 062678f2e6fe
model_pp0_ep116_shard0.safetensorsWeights4.2 GB 6a2903ea03b8
model_pp0_ep117_shard0.safetensorsWeights4.2 GB 87a2dfe3c8ec
model_pp0_ep118_shard0.safetensorsWeights4.2 GB 2a723bfe7f43
model_pp0_ep119_shard0.safetensorsWeights4.2 GB c1cbc2af036a
model_pp0_ep11_shard0.safetensorsWeights4.2 GB 5ec6c79fa658
model_pp0_ep120_shard0.safetensorsWeights4.2 GB 1a5927cdedb8
model_pp0_ep121_shard0.safetensorsWeights4.2 GB e131992cd440
model_pp0_ep122_shard0.safetensorsWeights4.2 GB 4fe86280f26b
model_pp0_ep123_shard0.safetensorsWeights4.2 GB b4c96ee7b894
model_pp0_ep124_shard0.safetensorsWeights4.2 GB 3fb9ef76464b
model_pp0_ep125_shard0.safetensorsWeights4.2 GB e21ac89a1588
model_pp0_ep126_shard0.safetensorsWeights4.2 GB 9337dd6c9c51
model_pp0_ep127_shard0.safetensorsWeights4.2 GB b478fb7f61d7
model_pp0_ep12_shard0.safetensorsWeights4.2 GB 0a114e99fd1e
model_pp0_ep13_shard0.safetensorsWeights4.2 GB 0f2c0187d966
model_pp0_ep14_shard0.safetensorsWeights4.2 GB c6566f910330
model_pp0_ep15_shard0.safetensorsWeights4.2 GB 336fe95d53df
model_pp0_ep16_shard0.safetensorsWeights4.2 GB ab149f7efe13
model_pp0_ep17_shard0.safetensorsWeights4.2 GB c273f0b57f8b
model_pp0_ep18_shard0.safetensorsWeights4.2 GB eac24bc4a640
model_pp0_ep19_shard0.safetensorsWeights4.2 GB b8af151d21d5
model_pp0_ep1_shard0.safetensorsWeights4.2 GB e267f008137b
model_pp0_ep20_shard0.safetensorsWeights4.2 GB 920b8f1bf257
model_pp0_ep21_shard0.safetensorsWeights4.2 GB 44dfe8476db6
model_pp0_ep22_shard0.safetensorsWeights4.2 GB 7e6ded1cf3f6
model_pp0_ep23_shard0.safetensorsWeights4.2 GB 5db120d005e6
model_pp0_ep24_shard0.safetensorsWeights4.2 GB 4bcc389dfb83
model_pp0_ep25_shard0.safetensorsWeights4.2 GB 8070ba6fabba
model_pp0_ep26_shard0.safetensorsWeights4.2 GB 379c98ff3c00
model_pp0_ep27_shard0.safetensorsWeights4.2 GB 979dee453997
model_pp0_ep28_shard0.safetensorsWeights4.2 GB 4c2ec47cd24f
model_pp0_ep29_shard0.safetensorsWeights4.2 GB 7da80d925a9f
model_pp0_ep2_shard0.safetensorsWeights4.2 GB 04db5974b8ec
model_pp0_ep30_shard0.safetensorsWeights4.2 GB 5f5599a35f9f
model_pp0_ep31_shard0.safetensorsWeights4.2 GB fbcd47f22bad
model_pp0_ep32_shard0.safetensorsWeights4.2 GB 68d63a7e59b9
model_pp0_ep33_shard0.safetensorsWeights4.2 GB b2cb840f19f2
model_pp0_ep34_shard0.safetensorsWeights4.2 GB 2cb7c5e0b597
model_pp0_ep35_shard0.safetensorsWeights4.2 GB dfa9ece9d5ab
model_pp0_ep36_shard0.safetensorsWeights4.2 GB 91249ea0740e
model_pp0_ep37_shard0.safetensorsWeights4.2 GB d3e24b54883a
model_pp0_ep38_shard0.safetensorsWeights4.2 GB a9192363aaea
model_pp0_ep39_shard0.safetensorsWeights4.2 GB 6843ec8872bc
model_pp0_ep3_shard0.safetensorsWeights4.2 GB 49b3bf51cc50
model_pp0_ep40_shard0.safetensorsWeights4.2 GB f6d29d9a0fcc
model_pp0_ep41_shard0.safetensorsWeights4.2 GB d220d567b191
model_pp0_ep42_shard0.safetensorsWeights4.2 GB cd207a813285
model_pp0_ep43_shard0.safetensorsWeights4.2 GB d047079d2aaf
model_pp0_ep44_shard0.safetensorsWeights4.2 GB 55e2dec8a549
model_pp0_ep45_shard0.safetensorsWeights4.2 GB 6017bad79034
model_pp0_ep46_shard0.safetensorsWeights4.2 GB 245f0358c8f2
model_pp0_ep47_shard0.safetensorsWeights4.2 GB 6f98a831c7b9
model_pp0_ep48_shard0.safetensorsWeights4.2 GB 2b5755d95b93
model_pp0_ep49_shard0.safetensorsWeights4.2 GB 0a54a726bb15
model_pp0_ep4_shard0.safetensorsWeights4.2 GB 58ac0432f6d0
model_pp0_ep50_shard0.safetensorsWeights4.2 GB 2e3a8eece5f0
model_pp0_ep51_shard0.safetensorsWeights4.2 GB 393738725c5a
model_pp0_ep52_shard0.safetensorsWeights4.2 GB 328c8447d45a
model_pp0_ep53_shard0.safetensorsWeights4.2 GB aace62146268
model_pp0_ep54_shard0.safetensorsWeights4.2 GB c030c5dd1248
model_pp0_ep55_shard0.safetensorsWeights4.2 GB 372a1dc790fa
model_pp0_ep56_shard0.safetensorsWeights4.2 GB b94507e90c1e
model_pp0_ep57_shard0.safetensorsWeights4.2 GB f317149ac0ac
model_pp0_ep58_shard0.safetensorsWeights4.2 GB 5d14b1efa101
model_pp0_ep59_shard0.safetensorsWeights4.2 GB 95e43b9f3a4d
model_pp0_ep5_shard0.safetensorsWeights4.2 GB a2afb6b46360
model_pp0_ep60_shard0.safetensorsWeights4.2 GB 1634e6c23b89
model_pp0_ep61_shard0.safetensorsWeights4.2 GB b7336d380377
model_pp0_ep62_shard0.safetensorsWeights4.2 GB 9d51a2539aef
model_pp0_ep63_shard0.safetensorsWeights4.2 GB 43cd3209a978
model_pp0_ep64_shard0.safetensorsWeights4.2 GB 80207bc42408
model_pp0_ep65_shard0.safetensorsWeights4.2 GB 25b67300895e
model_pp0_ep66_shard0.safetensorsWeights4.2 GB 953777d906e2
model_pp0_ep67_shard0.safetensorsWeights4.2 GB 92203f915388
model_pp0_ep68_shard0.safetensorsWeights4.2 GB 3bdb2c292f9c
model_pp0_ep69_shard0.safetensorsWeights4.2 GB 0d33f2b21b35
model_pp0_ep6_shard0.safetensorsWeights4.2 GB 2aefd5228285
model_pp0_ep70_shard0.safetensorsWeights4.2 GB 14db94064819
model_pp0_ep71_shard0.safetensorsWeights4.2 GB d680a315854e
model_pp0_ep72_shard0.safetensorsWeights4.2 GB d7e45c83608e
model_pp0_ep73_shard0.safetensorsWeights4.2 GB 36a19a06af82
model_pp0_ep74_shard0.safetensorsWeights4.2 GB 485728009a43
model_pp0_ep75_shard0.safetensorsWeights4.2 GB 54c1f2ce568d
model_pp0_ep76_shard0.safetensorsWeights4.2 GB 13fdf074bd99
model_pp0_ep77_shard0.safetensorsWeights4.2 GB 7d21e5436d92
model_pp0_ep78_shard0.safetensorsWeights4.2 GB 462e84439f32
model_pp0_ep79_shard0.safetensorsWeights4.2 GB fe82cba5169e
model_pp0_ep7_shard0.safetensorsWeights4.2 GB c7216645761d
model_pp0_ep80_shard0.safetensorsWeights4.2 GB 5e8102db99d7
model_pp0_ep81_shard0.safetensorsWeights4.2 GB 6121c4360850
model_pp0_ep82_shard0.safetensorsWeights4.2 GB facc722552eb
model_pp0_ep83_shard0.safetensorsWeights4.2 GB ef024960b281
model_pp0_ep84_shard0.safetensorsWeights4.2 GB 945186a9047c
model_pp0_ep85_shard0.safetensorsWeights4.2 GB c71846d5b5a3
model_pp0_ep86_shard0.safetensorsWeights4.2 GB d01073e4b10f
model_pp0_ep87_shard0.safetensorsWeights4.2 GB 98cffd997fca
model_pp0_ep88_shard0.safetensorsWeights4.2 GB 253f92111c26
model_pp0_ep89_shard0.safetensorsWeights4.2 GB 463ae82378ec
model_pp0_ep8_shard0.safetensorsWeights4.2 GB f5b2ae624427
model_pp0_ep90_shard0.safetensorsWeights4.2 GB 663a527ce7b9
model_pp0_ep91_shard0.safetensorsWeights4.2 GB 30505f129bf9
model_pp0_ep92_shard0.safetensorsWeights4.2 GB e1cf2487949f
model_pp0_ep93_shard0.safetensorsWeights4.2 GB c90af4cab01f
model_pp0_ep94_shard0.safetensorsWeights4.2 GB 2cff49450d67
model_pp0_ep95_shard0.safetensorsWeights4.2 GB 77d5fa0383d2
model_pp0_ep96_shard0.safetensorsWeights4.2 GB c9378becaae1
model_pp0_ep97_shard0.safetensorsWeights4.2 GB 5b35962e5778
model_pp0_ep98_shard0.safetensorsWeights4.2 GB 3b1bc2348cbd
model_pp0_ep99_shard0.safetensorsWeights4.2 GB dd2c6ca2e0ad
model_pp0_ep9_shard0.safetensorsWeights4.2 GB 5a1a3774c6ee
audio_tokenizer/config.jsonConfiguration1.2 KB —
audio_tokenizer/generation_config.jsonConfiguration149 B —
config.jsonConfiguration9.3 KB —
configuration_mimo_v2.pyConfiguration10.0 KB —
dflash/config.jsonConfiguration1.2 KB —
dflash/dflash.pyConfiguration14.3 KB —
dflash/model.safetensors.index.jsonConfiguration4.7 KB —
generation_config.jsonConfiguration195 B —
model.safetensors.index.jsonConfiguration15.2 MB e855ee9dd7ae
modeling_mimo_v2.pyConfiguration85.5 KB —
preprocessor_config.jsonConfiguration350 B —
README.mdDocumentation9.8 KB —
MiMo_V2_6_technical_report.pdfOther3.0 MB 7fe42601dc95
assets/architecture.pngOther405.4 KB d288768e1771
audio_tokenizer/chat_template.jinjaOther5.6 KB —
chat_template.jinjaOther3.9 KB —
.gitattributesRepository1.8 KB —
audio_tokenizer/tokenizer_config.jsonTokenizer6.1 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB ff15eb925890
tokenizer_config.jsonTokenizer12.2 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
mit
Access
Open weights, no gate
Download size
573.5 GB
Download from Xiaomi MiMo

Released by Xiaomi MiMo through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published573.5 GB
16-bit1048.2 GB
8-bit524.1 GB
4-bit262.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About MiMo-V2.6-Pro-RL

How much GPU memory does MiMo-V2.6-Pro-RL need?

About 1257.9 GB at 16-bit and 314.5 GB at 4-bit: the weights (524.1B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run MiMo-V2.6-Pro-RL on?

At 16-bit, 5x MI325X from $10.00 an hour; at 4-bit, 2x MI300X from $3.70 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use MiMo-V2.6-Pro-RL commercially?

Yes. MiMo-V2.6-Pro-RL is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is MiMo-V2.6-Pro-RL's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

MiMo-V2.6-Pro-RL

Zachary Howard

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn…

Open weights mit 524.1B parameters 1,048,576 tokens transformers

Model · Text generation

Qwen3-Coder-480B-A35B-Instruct-FP8

Qwen

Today, we're announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we're excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct. featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks, achieving results comparable to Claude Sonnet. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.…

Open weights apache-2.0 480.2B parameters 262,144 tokens transformers

Model · Text generation

Darwin-397B-ZTC

FINAL_Bench

the footprint, GPQA Diamond 93.43 %. And this model stops itself before it acts on an answer it is about to get wrong. Darwin is VIDRAFT's measurement-driven reasoning model family — roughly 20 official models, 400+ community derivatives, and a standing place among the top open models on GPQA. A large MoE model is made of hundreds of experts. Darwin V9 selects the experts that perform best across several high-performing models, transplants them onto a base backbone, and fuses them with trust-weighted evolutionary merging. Nothing is trained from scratch — proven capability is grafted on. That is why the same method holds across every model size. - Darwin V9 — evolutionary FFN/expert…

Open weights apache-2.0 403.6B parameters 262,144 tokens transformers

Model · Text generation

Darwin-397B-JGOS

FINAL_Bench

Darwin-397B-JGOS is the largest and highest-scoring member of the Darwin family. Built on Qwen 3.5 397B as the base, it transplants the FFN (expert) strengths of multiple high-performance models through the Darwin V9 platform, producing a 397B-parameter Mixture-of-Experts model with ~17B active parameters per token. It reaches 90.9 % on GPQA Diamond with pure greedy decoding (single sample) — surpassing Darwin-28B-REASON (89.39 %, achieved with the Darwin-DELPHI test-time engine) without using any test-time engine at all. This is the highest GPQA Diamond score in the Darwin family to date. Darwin is VIDRAFT's measuring-result-driven reasoning model family — approximately 20 official models…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Text generation

GLM-5.2-NVFP4

NVIDIA

The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM-5.2 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial or non-commercial use. GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG…

Open weights mit 381B parameters 1,048,576 tokens Model Optimizer

K2-Horizon-375B-A23B is the flagship of the K2-Horizon family: a sparse Mixture-of-Experts model that stores 375B parameters and runs 23B per token, with a 512K context window. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released. - Frontier-class agentic performance. On agentic tool use, terminal, and long-horizon workflow benchmarks it matches or beats open-weight MoE models up to 2.6× its size and is competitive with closed frontier models (see Benchmark Results). - 512K context. Native 524,288-token context from the midtraining stages onward. - Intermediate checkpoints. Intermediate checkpoints will be released so…

Open weights apache-2.0 379.2B parameters 524,288 tokens transformers