SAVRN
Search Contact SAVRN

Open-weight model · Text generation

MiMo-V2.6-Flash-RL

by Zachary Howard servantofares/MiMo-V2.6-Flash-RL

MiMo-V2.6-Flash-RL is an open-weight model for text generation from Zachary Howard, released under MIT License. It has 159.4B parameters and a 1,048,576-token context. At 16-bit it needs about 382.5 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 95.6 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index.

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series.

Parameters159.4B
Context1,048,576
Weights177.7 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve MiMo-V2.6-Flash-RL (159.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 318.7 GB 382.5 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
8-bit 159.4 GB 191.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 79.7 GB 95.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

MiMo-V2.6-Flash-RL on every accelerator the SAVRN Index prices, at every precision

Model Card

By Zachary Howard, published under mit, revision d5fca88077d4.

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Read Zachary Howard's full model card





Community
WeChat Group  |  Discord  |  Telegram  |  Reddit


MiMo-V2.6-Flash-RL

Scaling Reinforcement Learning Toward Self-Improvement

Technical Report

1. Introduction

MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include:

  • Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs.
  • You Only RL Once: One mixed RL run across coding, general agents, visual, and cybersecurity — not separate per-domain runs. Tasks and multiple harnesses are mixed in the same batch so capabilities reinforce each other and strategies transfer to harnesses never seen in training.
  • Scaling RL Compute: Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches — 1,568 prompts × 16 rollouts per step, billions of tokens per update.
  • Groupwise Agentic Grading (Self-Improvement Loop): Binary pass/fail cannot rank passing solutions, so the reward signal itself is scaled. An agentic grader compares rollouts within each group: Groupwise Reward Synthesis (GRS) builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes; Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.
  • Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
  • Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix), reusing histories from teacher trajectories and SFT demonstrations so decision points train without regenerating preceding turns — extending capabilities to hard-to-verify tasks.

Model Summary

  • Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
  • Context Length: 1M tokens
  • Modalities: Text, Image, Video, Audio
  • Vision Encoder: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
  • Audio Encoder: 308M AudioTokenizer + 127M audio patch encoder
  • Multi-Token Prediction (MTP): 5-layer speculative decoder

Figure 1. MiMo-V2.6 architecture.

2. Downloads

Model Download
MiMo-V2.6-Pro-RL HuggingFace· ModelScope(at release)
MiMo-V2.6-Flash-RL HuggingFace· ModelScope(at release)

3. Evaluation Results

Benchmark MiMo-V2.6 Pro MiMo-V2.6 Flash MiMo-V2.5 Pro Claude Opus 5 GPT-5.6 Sol Claude Fable 5
Code Agent
DeepSWE v1.1 71.9 67.9 19.0 74.0 73.0 70.0
ProgramBench 26.5 26.0 12.5 37.0 25.0 33.0
MiMo Code Bench 63.2 61.2 40.4 68.6 59.3 -
General Agent
AutomationBench v1.0.6 53.1 52.3 16.0 50.3 45.8 46.2
Toolathlon-Verified 76.9 73.6 49.1 80.6 74.9 77.9
GDPval-AA 2.1 1673 - 1107 1708 1588 1595
Agents’ Last Exam 31.6 27.6 13.2 31.6 30.8 25.7
Terminal Bench 4.0 34.9 28.8 1.5 49.0 39.9 42.4
Terminal Bench 2.1 89.9 87.6 65.2 89.1 88.8 84.3
OSWorld-Verified 82.0 80.8 - 83.4 83.0 86.0
JobBench 62.0 61.2 25.0 65.7 45.4 57.4
Cybersecurity
CyberGym 94.0 95.1 40.0 - - -
MiMo Cyber Bench 80.2 77.2 0.0 - - -
ExploitGym 17.8 6.0 0.2 22.1 30.3 28.4
ExploitBench 47.9 25.3 16.6 70.0 78.5 78.0
SEC Bench Pro 66.3 47.5 17.7 - 79.1 -
Visual Agent
MiMo VisualCoding 72.3 71.5 - 70.0 73.4 69.1

4. Model Architecture

LLM Backbone

Component MiMo-V2.6-Flash-RL
Layers (Total / SWA / GA) 48 / 39 / 9
Hidden Size 4096
SWA Heads (Q/KV) 64 / 8
GA Heads (Q/KV) 64 / 4
Head Dimensions (QK / V) 192 / 128
Sliding Window Size 128
Routed Experts (Total / Activated) 256 / 8
Max Context Length 1M
MTP / Speculative Decoder 5 SWA layers, window 1024

The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.

Vision Encoder (MiMo ViT)

Configuration Value
Layers (Total / SWA / GA) 28 / 24 / 4
Hidden Size 1280
Attention Heads (Q / KV) 32 / 8
Head Dimension 64
Patch Size (T × H × W) 2 × 16 × 16
Sliding Window (Left / Right) 64 / 64
Spatial Merge Size 2 × 2
Parameters 681M

Audio Encoders

AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).

Speculative Decoder

5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.

5. Deployment

For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.

SGLang

sglang serve \
  --trust-remote-code \
  --model-path XiaomiMiMo/MiMo-V2.6-Flash-RL \
  --tp 8 \
  --dp 2 \
  --enable-dp-attention \
  --enable-dp-lm-head \
  --mm-enable-dp-encoder \
  --mem-fraction-static 0.65 \
  --chunked-prefill-size 16384 \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --enable-multi-layer-eagle \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --host 0.0.0.0 \
  --port 30000

vLLM

Follow the vLLM MiMo-V2.5 recipe. Stable vLLM may lag; pre-built image: docker pull vllm/vllm-openai:mimov25-cu129.

vllm serve XiaomiMiMo/MiMo-V2.6-Flash-RL \
  --tensor-parallel-size 4 \
  --trust-remote-code \
  --gpu-memory-utilization 0.95 \
  --max-model-len auto \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --enable-auto-tool-choice \
  --generation-config vllm

Recommended sampling: temperature=1.0, top_p=0.95.

Also available in AI Studio, MiMo Code, Xiaomi MiMo Desktop, Xiaomi MiMo Open Platform API, and OpenRouter.

Citation

@misc{mimo2026v26flash,
  title={MiMo-V2.6-Flash-RL},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}},
}

Contact

For questions or feedback, reach us at [email protected] or join our community:

Configuration

Architecture
MiMoV2ForCausalLM
Context length (tokens)
1,048,576
Layers
48
Hidden size
4,096
Feed-forward size
16,384
Attention heads
64
Key/value heads
4
Head dimension
192
Vocabulary size
152,576
Routed experts
256
Experts active per token
8
Sliding window (tokens)
128
RoPE base
1e+07
Model type
mimo_v2
Quantization
fp8

Identity and Version

Repository
servantofares/MiMo-V2.6-Flash-RL
Publisher
Zachary Howard
Task
Text generation
Modality
Text
Library
transformers
Parameters
159.4B parameters
Languages
en, zh
Revision
d5fca88077d42261c8a9f5f48f5259789eccbcd6
First published
2026-09-22
Last updated
2026-09-22

Files and Weights

90 files, 177.8 GB in total. The weights are 68 files totalling 177.7 GB in pt, safetensors.

Weights68 files · 177.7 GB
Configuration11 files · 7.0 MB
Tokenizer5 files · 15.9 MB
Documentation1 file · 9.5 KB
Other4 files · 3.5 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
audio_tokenizer/model.safetensorsWeights1.9 GB 077033345d80
dflash/dflash_draft_model.safetensorsWeights2.9 GB 94d9c02c17e0
dflash/mask_embedding.ptWeights9.9 KB b35b379fe049
model_mtp.safetensorsWeights1.2 GB f8f78a78a79e
model_pp0_ep0_shard0.safetensorsWeights13.4 GB 88ff485461b9
model_pp0_ep10_shard0.safetensorsWeights2.5 GB da00ca1b796d
model_pp0_ep11_shard0.safetensorsWeights2.5 GB e71b67866924
model_pp0_ep12_shard0.safetensorsWeights2.5 GB 415a846ced38
model_pp0_ep13_shard0.safetensorsWeights2.5 GB cfdf506f2f41
model_pp0_ep14_shard0.safetensorsWeights2.5 GB 035ec01fe9ad
model_pp0_ep15_shard0.safetensorsWeights2.5 GB c68e1ccf79a6
model_pp0_ep16_shard0.safetensorsWeights2.5 GB ff707d5aeaec
model_pp0_ep17_shard0.safetensorsWeights2.5 GB 284187f225f7
model_pp0_ep18_shard0.safetensorsWeights2.5 GB f5a701a582e9
model_pp0_ep19_shard0.safetensorsWeights2.5 GB e9f48244d5d0
model_pp0_ep1_shard0.safetensorsWeights2.5 GB 22b6affecd08
model_pp0_ep20_shard0.safetensorsWeights2.5 GB a705193ebcd1
model_pp0_ep21_shard0.safetensorsWeights2.5 GB 00d5447c5189
model_pp0_ep22_shard0.safetensorsWeights2.5 GB 28a81aa95b2c
model_pp0_ep23_shard0.safetensorsWeights2.5 GB 9b26e23a99ee
model_pp0_ep24_shard0.safetensorsWeights2.5 GB 5a7d6d132b4e
model_pp0_ep25_shard0.safetensorsWeights2.5 GB e4381d38a657
model_pp0_ep26_shard0.safetensorsWeights2.5 GB e60972151139
model_pp0_ep27_shard0.safetensorsWeights2.5 GB 667920e72309
model_pp0_ep28_shard0.safetensorsWeights2.5 GB 1276d80fccbb
model_pp0_ep29_shard0.safetensorsWeights2.5 GB b14eeeff5a31
model_pp0_ep2_shard0.safetensorsWeights2.5 GB 09e0e60b70a4
model_pp0_ep30_shard0.safetensorsWeights2.5 GB 71a1cafd2049
model_pp0_ep31_shard0.safetensorsWeights2.5 GB 0d63c0b26ffc
model_pp0_ep32_shard0.safetensorsWeights2.5 GB 3227c1161ea3
model_pp0_ep33_shard0.safetensorsWeights2.5 GB 454ea014bf04
model_pp0_ep34_shard0.safetensorsWeights2.5 GB ca87893994c7
model_pp0_ep35_shard0.safetensorsWeights2.5 GB d32d42276805
model_pp0_ep36_shard0.safetensorsWeights2.5 GB f683b4cc856c
model_pp0_ep37_shard0.safetensorsWeights2.5 GB a1ec318aceec
model_pp0_ep38_shard0.safetensorsWeights2.5 GB dcbdeabcb804
model_pp0_ep39_shard0.safetensorsWeights2.5 GB e2355ec4c04a
model_pp0_ep3_shard0.safetensorsWeights2.5 GB 040c1b7c7819
model_pp0_ep40_shard0.safetensorsWeights2.5 GB 18fb96b14541
model_pp0_ep41_shard0.safetensorsWeights2.5 GB 040930a123bc
model_pp0_ep42_shard0.safetensorsWeights2.5 GB 2d6946e36278
model_pp0_ep43_shard0.safetensorsWeights2.5 GB f7bc2229a9d3
model_pp0_ep44_shard0.safetensorsWeights2.5 GB 2cb175e47f44
model_pp0_ep45_shard0.safetensorsWeights2.5 GB 24ddfb9ef70c
model_pp0_ep46_shard0.safetensorsWeights2.5 GB 1c911acaf741
model_pp0_ep47_shard0.safetensorsWeights2.5 GB 6a3454f4223e
model_pp0_ep48_shard0.safetensorsWeights2.5 GB 6a72d2de5e67
model_pp0_ep49_shard0.safetensorsWeights2.5 GB 34012530cff5
model_pp0_ep4_shard0.safetensorsWeights2.5 GB f166afbd8d82
model_pp0_ep50_shard0.safetensorsWeights2.5 GB c3d219144d63
model_pp0_ep51_shard0.safetensorsWeights2.5 GB e378611cc5a1
model_pp0_ep52_shard0.safetensorsWeights2.5 GB bf3628acbd82
model_pp0_ep53_shard0.safetensorsWeights2.5 GB b1847ad14ae3
model_pp0_ep54_shard0.safetensorsWeights2.5 GB 26755fdb7dd5
model_pp0_ep55_shard0.safetensorsWeights2.5 GB 560164ba11d0
model_pp0_ep56_shard0.safetensorsWeights2.5 GB 0e7a659c8453
model_pp0_ep57_shard0.safetensorsWeights2.5 GB 401978e73cf9
model_pp0_ep58_shard0.safetensorsWeights2.5 GB e2cc332b6f7c
model_pp0_ep59_shard0.safetensorsWeights2.5 GB 442203f3ba29
model_pp0_ep5_shard0.safetensorsWeights2.5 GB bf32dcbbc1bc
model_pp0_ep60_shard0.safetensorsWeights2.5 GB 60376b2bf894
model_pp0_ep61_shard0.safetensorsWeights2.5 GB 18069789870b
model_pp0_ep62_shard0.safetensorsWeights2.5 GB deb270973083
model_pp0_ep63_shard0.safetensorsWeights2.5 GB 4c7d30662358
model_pp0_ep6_shard0.safetensorsWeights2.5 GB 078ba6e5316b
model_pp0_ep7_shard0.safetensorsWeights2.5 GB d835940a8c46
model_pp0_ep8_shard0.safetensorsWeights2.5 GB 8b85464b513d
model_pp0_ep9_shard0.safetensorsWeights2.5 GB c4e94cfa075d
audio_tokenizer/config.jsonConfiguration1.2 KB —
audio_tokenizer/generation_config.jsonConfiguration149 B —
config.jsonConfiguration8.1 KB —
configuration_mimo_v2.pyConfiguration10.0 KB —
dflash/config.jsonConfiguration1.2 KB —
dflash/dflash.pyConfiguration14.3 KB —
dflash/model.safetensors.index.jsonConfiguration4.7 KB —
generation_config.jsonConfiguration195 B —
model.safetensors.index.jsonConfiguration6.9 MB —
modeling_mimo_v2.pyConfiguration85.5 KB —
preprocessor_config.jsonConfiguration350 B —
README.mdDocumentation9.5 KB —
MiMo_V2_6_technical_report.pdfOther3.0 MB 7fe42601dc95
assets/architecture.pngOther405.4 KB d288768e1771
audio_tokenizer/chat_template.jinjaOther5.6 KB —
chat_template.jinjaOther3.9 KB —
.gitattributesRepository1.7 KB —
audio_tokenizer/tokenizer_config.jsonTokenizer6.1 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB ff15eb925890
tokenizer_config.jsonTokenizer12.2 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
mit
Access
Open weights, no gate
Download size
177.7 GB
Download from Zachary Howard

Released by Zachary Howard through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published177.7 GB
16-bit318.7 GB
8-bit159.4 GB
4-bit79.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About MiMo-V2.6-Flash-RL

How much GPU memory does MiMo-V2.6-Flash-RL need?

About 382.5 GB at 16-bit and 95.6 GB at 4-bit: the weights (159.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run MiMo-V2.6-Flash-RL on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use MiMo-V2.6-Flash-RL commercially?

Yes. MiMo-V2.6-Flash-RL is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is MiMo-V2.6-Flash-RL's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

MiMo-V2.6-Flash-RL

Xiaomi MiMo

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers

Model · Text generation

Hy3-Razor-154B-A18B-E96of192

Mingyang Song

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…

Open weights apache-2.0 153.8B parameters 262,144 tokens transformers

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

Model · Text generation

Vinci-Cyber-123B-1.0

SimpleDirect

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…

Open weights other 125B parameters 262,144 tokens transformers

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Open weights other 124B parameters 262,144 tokens vllm

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers