Work in progress — model card published ahead of the weights, 2026-09-22. Original-checkpoint text importance-matrix collection is running. Multimodal input alignment, media calibration, mixed quantization, and final quality checks remain in progress. No release weights or final benchmark results are published yet.
One planned role-aware variant: routed expert gate/up at IQ2_XXS, routed expert down at IQ2_XS, attention/dense/embedding/head/MTP matrices at Q8_0, and multimodal encoder/projector matrices at BF16. Router, normalization, and control tensors retain F32.
Estimated weight storage: approximately 93.09 GB / 86.69 GiB, including the language model with three embedded MTP blocks, the multimodal projector, and the separate DFlash draft model. This is a build estimate, not measured device residency or a confirmed DGX Spark memory budget.
The target is preservation of text, image, audio, video, and joint audiovisual inputs. End-to-end validation is pending. DGX Spark qualification and MiMo support in ds4-dfm-rs are a separate follow-up; this card does not claim that the current runtime serves these artifacts.
Support my work
I work on making large language models practical on constrained hardware through mixed quantization, inference optimization, and serving experiments. Contributions help cover calibration, GPU compute, storage, and testing so these results can be published openly.
This is a mixed-precision conversion of XiaomiMiMo/MiMo-V2.6-Flash-RL, pinned to revision 3b38d063180c3e4aed9691fdc735f3d10b266ee4.
Why this model needs an asymmetric layout
The source stores routed expert weights as packed MXFP4. Two logical weights occupy each packed byte, so counting stored tensor elements can produce a roughly 159B figure. The expanded language trunk contains approximately 308.78B logical parameters; the root checkpoint including embedded MTP and media components contains approximately 310.76B, excluding the separate DFlash package. The packed representation does not make this a 159B logical model.
Routed experts account for 302.80B parameters. They dominate the storage budget, so the recipe concentrates compression there and gives their output projections a higher tier. Shared attention and dense paths, the output head, and multimodal components receive substantially more precision.
Variant and availability
| Component |
Precision |
Estimated size, decimal GB |
Status |
| Main model, including three embedded MTP blocks |
IQ2_XXS / IQ2_XS / Q8_0 / F32 |
88.77 |
Calibration running; final quantization pending |
| Multimodal encoder/projector |
BF16 / F32 |
2.75 |
Converted locally; input parity and validation pending |
| Separate DFlash draft model |
Q8_0 / F32 |
1.56 |
Converted locally; validation pending |
| All components |
Mixed |
93.09 |
Estimate; weights not yet released |
The main-plus-media estimate without the optional separate DFlash model is 91.52 GB / 85.23 GiB. Exact release filenames, shard counts, download commands, and checksums will be added after the final artifacts pass validation.
Quantization targets
| Model region |
Target |
Reason |
Routed expert gate and up |
IQ2_XXS |
Largest share of the parameter budget |
Routed expert down |
IQ2_XS |
Higher precision at the return to the residual stream |
| Attention projections and dense FFN matrices |
Q8_0 |
Shared computation on every token |
| Token embedding and output head |
Q8_0 |
Preserve input and logit precision |
| Three embedded MTP blocks, eligible matrices |
Q8_0 |
Preserve checkpoint predictors for later runtime integration |
| Routers, norms, attention sinks, and control tensors |
F32 |
Preserve routing and numerical control |
| Multimodal encoder/projector matrices |
BF16 |
Preserve media representation precision |
| Separate DFlash draft matrices |
Q8_0 |
Optional draft package, separate from embedded MTP |
Final per-tensor inventories will accompany the weights. BF16/F32 in this table describes the output storage format; source tensors already stored in FP8 are expanded from that source precision.
Importance matrix and calibration
Calibration uses the original checkpoint representation: expert MXFP4 values are repacked into GGUF without an additional quantization step, while the source FP8 dense tensors are expanded to BF16. This reference is not an original full-BF16 checkpoint and is not calibrated from the final IQ2 artifact.
The corpus starts from Baekpica/Inkling-Small-Multimodal-Calibration, with media recovered from the source datasets and prompts tokenized using MiMo's own tokenizer. Inkling token IDs and embeddings are not reused. Video clips are added from FineVideo.
| Prepared subset |
Records |
Current state |
| Text reasoning and code/tool text |
667 |
Text imatrix collection running; 640 × 1,024-token chunks requested |
| Chart/document images |
486 |
Media restored; calibration pending |
| Audio |
309 |
Media restored; calibration pending |
| FineVideo clips |
79 |
Timestamp-aligned visual clips prepared; calibration pending |
| FineVideo holdout |
9 |
Separated by source video; not part of calibration |
The prepared visual clips are currently silent. Original audio recovery and joint audiovisual input alignment remain required before claiming audiovisual calibration. Routed-expert observation coverage will be audited before the final importance matrix is accepted. Prepared counts are not completed calibration counts.
Multimodal input contract
Input processing is being aligned with the SGLang MiMo implementation, alongside the pinned checkpoint configuration and tokenizer.
| Input |
Required handling |
Validation state |
| Text |
MiMo tokenizer, chat template, and special tokens |
Reference text decode passed |
| Image |
Production normalization, spatial patch/merge order, vision boundary tokens |
Native encoder smoke passed; parity checks pending |
| Audio |
Original audio tokenizer, local encoder/projector, audio boundary tokens |
Alignment and calibration pending |
| Video |
Two-frame temporal patches, MiMo video boundaries, MM:SS timestamps |
Adapter preparation in progress |
| Video with audio |
Source timing and native audiovisual token layout |
Preparation and validation pending |
Generic video ingestion is insufficient here: using independent image frames or generic timestamps changes the model's input contract. Media imatrix collection must consume the correctly prepared original-model embeddings.
Memory metrics
| Metric |
Bytes |
| Routed expert gate/up target payload |
52,042,924,032 |
| Routed expert down target payload |
29,175,578,624 |
| Dense matrices including MTP target payload |
7,354,253,312 |
| Router/norm/control target payload |
199,026,176 |
| Main target tensor payload |
88,771,782,144 |
| Local multimodal projector file |
2,748,509,792 |
| Separate DFlash estimated size |
1,559,730,688 |
Main-model figures are block-format payload calculations. Final GGUF headers, alignment, tokenizer metadata, and any additional runtime assets are separate. Device usage also includes KV cache, media activations, allocator overhead, workspaces, and the operating system. No context length or concurrency level is yet qualified on DGX Spark.
The serving plan leaves media encoders and the separate drafter eligible for selective loading and uses bounded media-processing buffers. Those runtime strategies still require implementation and measurement; retaining all modality assets does not imply loading every optional component simultaneously.
Conversion and verification
| Check |
Current evidence |
| Source download |
All 90 source repository files passed Hub checksum verification |
| MXFP4 repacking and TP=4 QKV ordering |
Local numerical checks passed |
| Routed-expert source audit |
All 36,096 expert matrices sampled across 108,257 deterministic rows; repacked bytes matched |
| Original-representation reference |
Four GGUF shards generated locally |
| Reference text smoke |
GPU load and arithmetic decode passed |
| Text imatrix |
Running; intermediate checkpoints saved |
| Multimodal imatrix and coverage |
Pending |
| Final mixed tensor inventory and checksums |
Pending |
| Final text/media quality comparison |
Pending |
| Embedded MTP and DFlash runtime validation |
Pending |
| DGX Spark / ds4-dfm-rs qualification |
Deferred to the Spark integration stage |
The sampled expert audit is not a whole-file byte comparison. A successful reference smoke is not a quality benchmark for the final mixed model. No throughput, accuracy, or quality-retention claim is made for the unreleased variant.
Reproduction and release contents
The release is planned to include split mixed GGUF weights, the BF16 multimodal projector, the separate Q8_0 DFlash package, source revision and tokenizer provenance, the per-tensor recipe, calibration manifests and coverage evidence, conversion scripts, checksums, and validation reports. Additional audio assets needed by the native input path will be inventoried before release.
A full Q8_0 or BF16 language baseline is not required by this recipe. Intermediate models are produced only when needed for calibration or validation.
Chat template
The upstream chat_template.jinja is published alongside this card, copied byte for byte from the pinned source revision. Use it with MiMo’s own tokenizer and special-token mapping. Media preprocessing and embedding insertion remain part of the runtime input contract; the template alone does not implement them.
Runtime support
The current work covers calibration, mixed-quant construction, validation, and a reproducible handoff. MiMo integration into ds4-dfm-rs and DGX Spark serving tests follow separately. Runnable serving commands will be published after the corresponding runtime path is tested.
License
MIT, inherited from the pinned upstream model. Calibration datasets retain their respective licenses and access conditions.