SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF

by Sunghoon Baek Baekpica/MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF

MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF is an open-weight model for image and text to text from Sunghoon Baek, released under MIT License. Its published files total 17.0 KB.

I work on making large language models practical on constrained hardware through mixed quantization, inference optimization, and serving experiments.

Parameters—
Context—
Weights17.0 KB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By Sunghoon Baek, published under mit, revision d4f44d8b13b5.

I work on making large language models practical on constrained hardware through mixed quantization, inference optimization, and serving experiments. Contributions help cover calibration, GPU compute, storage, and testing so these results can be published openly. This is a mixed-precision conversion of XiaomiMiMo/MiMo-V2.6-Flash-RL, pinned to revision 3b38d063180c3e4aed9691fdc735f3d10b266ee4. The source stores routed expert weights as packed MXFP4. Two logical weights occupy each packed byte, so counting stored tensor elements can produce a roughly 159B figure. The expanded language trunk contains approximately 308.78B logical parameters; the root checkpoint including embedded MTP and media…

Read Sunghoon Baek's full model card

Work in progress — model card published ahead of the weights, 2026-09-22. Original-checkpoint text importance-matrix collection is running. Multimodal input alignment, media calibration, mixed quantization, and final quality checks remain in progress. No release weights or final benchmark results are published yet.

One planned role-aware variant: routed expert gate/up at IQ2_XXS, routed expert down at IQ2_XS, attention/dense/embedding/head/MTP matrices at Q8_0, and multimodal encoder/projector matrices at BF16. Router, normalization, and control tensors retain F32.

Estimated weight storage: approximately 93.09 GB / 86.69 GiB, including the language model with three embedded MTP blocks, the multimodal projector, and the separate DFlash draft model. This is a build estimate, not measured device residency or a confirmed DGX Spark memory budget.

The target is preservation of text, image, audio, video, and joint audiovisual inputs. End-to-end validation is pending. DGX Spark qualification and MiMo support in ds4-dfm-rs are a separate follow-up; this card does not claim that the current runtime serves these artifacts.

Support my work

I work on making large language models practical on constrained hardware through mixed quantization, inference optimization, and serving experiments. Contributions help cover calibration, GPU compute, storage, and testing so these results can be published openly.

This is a mixed-precision conversion of XiaomiMiMo/MiMo-V2.6-Flash-RL, pinned to revision 3b38d063180c3e4aed9691fdc735f3d10b266ee4.

Why this model needs an asymmetric layout

The source stores routed expert weights as packed MXFP4. Two logical weights occupy each packed byte, so counting stored tensor elements can produce a roughly 159B figure. The expanded language trunk contains approximately 308.78B logical parameters; the root checkpoint including embedded MTP and media components contains approximately 310.76B, excluding the separate DFlash package. The packed representation does not make this a 159B logical model.

Routed experts account for 302.80B parameters. They dominate the storage budget, so the recipe concentrates compression there and gives their output projections a higher tier. Shared attention and dense paths, the output head, and multimodal components receive substantially more precision.

Variant and availability

Component Precision Estimated size, decimal GB Status
Main model, including three embedded MTP blocks IQ2_XXS / IQ2_XS / Q8_0 / F32 88.77 Calibration running; final quantization pending
Multimodal encoder/projector BF16 / F32 2.75 Converted locally; input parity and validation pending
Separate DFlash draft model Q8_0 / F32 1.56 Converted locally; validation pending
All components Mixed 93.09 Estimate; weights not yet released

The main-plus-media estimate without the optional separate DFlash model is 91.52 GB / 85.23 GiB. Exact release filenames, shard counts, download commands, and checksums will be added after the final artifacts pass validation.

Quantization targets

Model region Target Reason
Routed expert gate and up IQ2_XXS Largest share of the parameter budget
Routed expert down IQ2_XS Higher precision at the return to the residual stream
Attention projections and dense FFN matrices Q8_0 Shared computation on every token
Token embedding and output head Q8_0 Preserve input and logit precision
Three embedded MTP blocks, eligible matrices Q8_0 Preserve checkpoint predictors for later runtime integration
Routers, norms, attention sinks, and control tensors F32 Preserve routing and numerical control
Multimodal encoder/projector matrices BF16 Preserve media representation precision
Separate DFlash draft matrices Q8_0 Optional draft package, separate from embedded MTP

Final per-tensor inventories will accompany the weights. BF16/F32 in this table describes the output storage format; source tensors already stored in FP8 are expanded from that source precision.

Importance matrix and calibration

Calibration uses the original checkpoint representation: expert MXFP4 values are repacked into GGUF without an additional quantization step, while the source FP8 dense tensors are expanded to BF16. This reference is not an original full-BF16 checkpoint and is not calibrated from the final IQ2 artifact.

The corpus starts from Baekpica/Inkling-Small-Multimodal-Calibration, with media recovered from the source datasets and prompts tokenized using MiMo's own tokenizer. Inkling token IDs and embeddings are not reused. Video clips are added from FineVideo.

Prepared subset Records Current state
Text reasoning and code/tool text 667 Text imatrix collection running; 640 × 1,024-token chunks requested
Chart/document images 486 Media restored; calibration pending
Audio 309 Media restored; calibration pending
FineVideo clips 79 Timestamp-aligned visual clips prepared; calibration pending
FineVideo holdout 9 Separated by source video; not part of calibration

The prepared visual clips are currently silent. Original audio recovery and joint audiovisual input alignment remain required before claiming audiovisual calibration. Routed-expert observation coverage will be audited before the final importance matrix is accepted. Prepared counts are not completed calibration counts.

Multimodal input contract

Input processing is being aligned with the SGLang MiMo implementation, alongside the pinned checkpoint configuration and tokenizer.

Input Required handling Validation state
Text MiMo tokenizer, chat template, and special tokens Reference text decode passed
Image Production normalization, spatial patch/merge order, vision boundary tokens Native encoder smoke passed; parity checks pending
Audio Original audio tokenizer, local encoder/projector, audio boundary tokens Alignment and calibration pending
Video Two-frame temporal patches, MiMo video boundaries, MM:SS timestamps Adapter preparation in progress
Video with audio Source timing and native audiovisual token layout Preparation and validation pending

Generic video ingestion is insufficient here: using independent image frames or generic timestamps changes the model's input contract. Media imatrix collection must consume the correctly prepared original-model embeddings.

Memory metrics

Metric Bytes
Routed expert gate/up target payload 52,042,924,032
Routed expert down target payload 29,175,578,624
Dense matrices including MTP target payload 7,354,253,312
Router/norm/control target payload 199,026,176
Main target tensor payload 88,771,782,144
Local multimodal projector file 2,748,509,792
Separate DFlash estimated size 1,559,730,688

Main-model figures are block-format payload calculations. Final GGUF headers, alignment, tokenizer metadata, and any additional runtime assets are separate. Device usage also includes KV cache, media activations, allocator overhead, workspaces, and the operating system. No context length or concurrency level is yet qualified on DGX Spark.

The serving plan leaves media encoders and the separate drafter eligible for selective loading and uses bounded media-processing buffers. Those runtime strategies still require implementation and measurement; retaining all modality assets does not imply loading every optional component simultaneously.

Conversion and verification

Check Current evidence
Source download All 90 source repository files passed Hub checksum verification
MXFP4 repacking and TP=4 QKV ordering Local numerical checks passed
Routed-expert source audit All 36,096 expert matrices sampled across 108,257 deterministic rows; repacked bytes matched
Original-representation reference Four GGUF shards generated locally
Reference text smoke GPU load and arithmetic decode passed
Text imatrix Running; intermediate checkpoints saved
Multimodal imatrix and coverage Pending
Final mixed tensor inventory and checksums Pending
Final text/media quality comparison Pending
Embedded MTP and DFlash runtime validation Pending
DGX Spark / ds4-dfm-rs qualification Deferred to the Spark integration stage

The sampled expert audit is not a whole-file byte comparison. A successful reference smoke is not a quality benchmark for the final mixed model. No throughput, accuracy, or quality-retention claim is made for the unreleased variant.

Reproduction and release contents

The release is planned to include split mixed GGUF weights, the BF16 multimodal projector, the separate Q8_0 DFlash package, source revision and tokenizer provenance, the per-tensor recipe, calibration manifests and coverage evidence, conversion scripts, checksums, and validation reports. Additional audio assets needed by the native input path will be inventoried before release.

A full Q8_0 or BF16 language baseline is not required by this recipe. Intermediate models are produced only when needed for calibration or validation.

Chat template

The upstream chat_template.jinja is published alongside this card, copied byte for byte from the pinned source revision. Use it with MiMo’s own tokenizer and special-token mapping. Media preprocessing and embedding insertion remain part of the runtime input contract; the template alone does not implement them.

Runtime support

The current work covers calibration, mixed-quant construction, validation, and a reproducible handoff. MiMo integration into ds4-dfm-rs and DGX Spark serving tests follow separately. Runnable serving commands will be published after the corresponding runtime path is tested.

License

MIT, inherited from the pinned upstream model. Calibration datasets retain their respective licenses and access conditions.

Identity and Version

Repository
Baekpica/MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF
Publisher
Sunghoon Baek
Task
Image and text to text
Modality
Image and text
Library
gguf
Parameters
Not stated by the source
Languages
en, zh
Revision
d4f44d8b13b583daf72e2cbdee27db1c9714b361
First published
2026-09-22
Last updated
2026-09-22

Files and Weights

3 files, 17.0 KB in total.

Documentation1 file · 11.6 KB
Other1 file · 3.9 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
README.mdDocumentation11.6 KB —
chat_template.jinjaOther3.9 KB —
.gitattributesRepository1.5 KB —

License and Download

License
mit
Access
Open weights, no gate
Download from Sunghoon Baek

Released by Sunghoon Baek through its official repository on Hugging Face. Read the license.

Built From

Questions About MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF

Can I use MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF commercially?

Yes. MiMo-V2.6-Flash-RL-Mixed-Quant-GGUF is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF (Some of the weights are converted from PTQ1 to Q2K or Q3K.). This is just a test/validation. The ternary hybrid-attention kernels live in the PrismML-Eng/llama.cpp fork.…

Open weights apache-2.0 transformers

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. (there is also a extra robust "tools" version too.) A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. Extreme intelligence in a small package. Jaw dropping performance. Superior instruction following. A multi-stage and multi-model…

Open weights apache-2.0

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Non-uniform GGUF quantizations produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-27B at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. Both methods were developed at the…

Open weights apache-2.0 gguf