Open-weight model
MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw
by ramGPT ramgpt/MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw
MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw is an open-weight model from ramGPT, released under MIT License. It has 3.4B parameters and a 262,144-token context. At 16-bit it needs about 8.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. Validated locally on an NVIDIA GeForce RTX 4090. The throughput figures are short local smoke measurements, not standardized cross-system benchmarks.
Runs On
What it takes to serve MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw (3.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 6.8 GB | 8.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 3.4 GB | 4.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 1.7 GB | 2.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 6, 2026.
Model Card
By ramGPT, published under mit, revision 81328546105e.
EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. Validated locally on an NVIDIA GeForce RTX 4090. The throughput figures are short local smoke measurements, not standardized cross-system benchmarks. Tool-call validation produced a structured getweather call for Toronto. Vision validation used a generated test image containing a large red square; the model correctly returned red. The source weights were not modified. Two local metadata normalizations were required for the current ExLlamaV3 conversion path: 1. The Qwen3.5 processor metadata declared Qwen2VLImageProcessor; the local conversion copy was normalized to Qwen2VLImageProcessorFast to match the ExLlamaV3 Qwen3.5…
Read ramGPT's full model card
EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B.
Build provenance
- Source revision: 2367e865d009c13ac81713a2878291d33ab28177
- ExLlamaV3: 1.5.2+cu128.torch2.10.0
- PyTorch: 2.10.0+cu128
- Target bitrate: 4.00 bpw
- Architecture: Qwen3_5ForConditionalGeneration
- Source license: MIT
- Artifact size: approximately 6.4 GiB
Validation
Validated locally on an NVIDIA GeForce RTX 4090.
| Check | Result |
|---|---|
| Direct ExLlamaV3 load + generation | PASS |
| Direct decode smoke | 118.009 tok/s |
| TabbyAPI model load | PASS |
| OpenAI-compatible chat completion | PASS |
| TabbyAPI chat decode smoke | 147.64 tok/s |
| Streaming | PASS |
| Qwen3.5 tool-call parsing | PASS |
| Vision | PASS |
| Two simultaneous API requests | PASS |
| TabbyAPI configured context | 8192 tokens |
The throughput figures are short local smoke measurements, not standardized cross-system benchmarks.
Tool-call validation produced a structured get_weather call for Toronto.
Vision validation used a generated test image containing a large red square; the model correctly returned red.
Source metadata compatibility notes
The source weights were not modified. Two local metadata normalizations were required for the current ExLlamaV3 conversion path:
- The Qwen3.5 processor metadata declared Qwen2VLImageProcessor; the local conversion copy was normalized to Qwen2VLImageProcessorFast to match the ExLlamaV3 Qwen3.5 loader expectation.
- The source config declared text_config.mtp_num_hidden_layers = 1, while the source safetensors index contained no MTP/NextN tensors. The local conversion copy therefore disabled the absent MTP side model.
These changes affect conversion metadata only and are recorded in the local build provenance.
TabbyAPI configuration used
model:
backend: exllamav3
max_seq_len: 8192
cache_size: 8192
gpu_split_auto: false
gpu_split: [23, 0]
vision: true
reasoning: true
reasoning_start_token: "<think>"
reasoning_end_token: "</think>"
tool_format: qwen3_5
The validation environment intentionally excluded the RTX 3060 from model placement so the reported 4090 measurements are not distorted by heterogeneous multi-GPU splitting.
Notes
This is an independent community quantization of the XiaomiMiMo upstream model. Refer to the upstream model card for intended use, limitations, and license terms.
RTX 4090 context scaling
TabbyAPI, RTX 4090 only, EXL3 4.0 bpw, FP16 cache, vision enabled, max batch size 1. Each context test used a unique long prompt containing a hidden needle near the end; the response had to begin with the exact needle code.
| Target | Actual prompt | Prefill tok/s | Decode tok/s | Total time | Needle | Observed GPU0 memory |
|---|---|---|---|---|---|---|
| 1K | 1,038 | 4,943 | 145.2 | 0.91s | PASS | 8039 MiB |
| 4K | 4,104 | 8,922 | 144.1 | 1.09s | PASS | 8175 MiB |
| 8K | 8,199 | 9,212 | 140.5 | 1.45s | PASS | 8191 MiB |
| 16K | 16,392 | 8,957 | 137.2 | 2.37s | PASS | 8199 MiB |
| 32K | 32,541 | 4,828 | 127.2 | 7.31s | PASS | 8201 MiB |
| 64K | 64,587 | 6,577 | 111.3 | 10.48s | PASS | 8203 MiB |
All six needle checks passed through about 64K prompt tokens. Decode throughput fell gradually from about 145 tok/s at 1K to 111 tok/s at 64K.
These are single-pass local smoke measurements, not standardized benchmark scores. Prefill throughput can move non-monotonically with kernel compilation, autotuning, cache state and prompt shape. The GPU-memory column is total observed GPU0 usage during the run, not an isolated model allocation.
Raw results: benchmarks/rtx4090-tabbyapi-context.json.
Long-generation / repetition stability
I also ran six long-generation samples through TabbyAPI:
- 3 runs with thinking disabled.
- 3 runs with thinking enabled and a 512-token reasoning budget.
- Up to 1,024 generated tokens per run.
- Observed decode throughput stayed around 145-146 tok/s.
- 0/6 runs triggered the repeated n-gram loop heuristic.
- The largest repeated 12-word-window count was 2.
- Repeated 4-gram ratios ranged from 0.0079 to 0.0381.
- 4/6 runs reached the 1,024-token ceiling, so this should not be interpreted as an EOS-behavior test.
Some outputs repeated short Markdown-formatting lines; those are retained in the raw data but are not counted as semantic repetition loops unless longer n-grams repeat as well.
This is a small stability smoke test, not proof that repetition loops cannot occur with other prompts or sampling settings.
Raw results: benchmarks/rtx4090-loop-stability.json.
Hardware comparison: RTX 4090 vs RTX 3060 vs mixed split
The same EXL3 artifact was tested three ways. These are local TabbyAPI smoke measurements, not standardized benchmark scores.
| Placement | ~1K decode tok/s | ~8K decode tok/s | ~16K decode tok/s | Short generation tok/s |
|---|---|---|---|---|
| RTX 4090 only | 145.2 | 140.5 | 137.3 | ~145-148 |
| RTX 4090 + RTX 3060 layer split | 106.2 | 102.4 | 100.7 | 108.1 |
| RTX 3060 12GB only | 53.5 | 49.0 | 50.3 | 55.8 |
RTX 3060 12GB only
The model ran entirely on the RTX 3060 with no CPU model offload and no model layers placed on the RTX 4090. Vision was disabled for this accessibility test.
- 1K prompt: 53.5 tok/s, needle PASS
- 8K prompt: 49.0 tok/s, needle PASS
- ~16K prompt: 50.3 tok/s, needle PASS
- Short normal generation: 55.8 tok/s
- Repeated 4-gram ratio: 0.0
- Max repeated 12-gram count: 1
- Obvious repetition loop detected: no
Observed RTX 3060 memory was about 6.4-6.6 GiB during these runs. That is total observed GPU memory, not a model-only allocation measurement.
RTX 4090 + RTX 3060 mixed layer split
The mixed test used a conventional layer split with gpu_split [4, 8], not tensor parallelism.
- 1K prompt: 106.2 tok/s
- 8K prompt: 102.4 tok/s
- ~16K prompt: 100.7 tok/s
- Short normal generation: 108.1 tok/s
- Needle retrieval passed at all three context sizes.
- No obvious repetition loop was detected in the short generation smoke.
After load, total observed GPU memory was approximately 4.74 GiB on the RTX 4090 and 2.65 GiB on the RTX 3060.
For this 9B model, the mixed layer split is slower than a 4090-only placement because the model already fits comfortably on the faster card. The mixed result is still useful as a proof that heterogeneous layer splitting works and remains usable when a future model no longer fits on one GPU.
Raw results:
- benchmarks/rtx3060-tabbyapi.json
- benchmarks/rtx4090-rtx3060-mixed.json
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 32
- Hidden size
- 4,096
- Feed-forward size
- 12,288
- Attention heads
- 16
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
- Quantization
- exl3
Identity and Version
- Repository
- ramgpt/MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw
- Publisher
- ramGPT
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 3.4B parameters
- Languages
- Not stated by the source
- Revision
- 81328546105e5906af880578aecade33e2292383
- First published
- 2026-09-30
- Last updated
- 2026-09-30
Files and Weights
20 files, 6.9 GB in total. The weights are 2 files totalling 6.8 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00002.safetensors | Weights | 4.2 GB | 8262ad181237 |
| model-00002-of-00002.safetensors | Weights | 2.6 GB | 819f9002d461 |
| benchmarks/rtx3060-tabbyapi.json | Configuration | 2.9 KB | — |
| benchmarks/rtx4090-loop-stability.json | Configuration | 2.9 KB | — |
| benchmarks/rtx4090-rtx3060-mixed.json | Configuration | 2.7 KB | — |
| benchmarks/rtx4090-tabbyapi-context.json | Configuration | 2.9 KB | — |
| config.json | Configuration | 3.6 KB | — |
| generation_config.json | Configuration | 126 B | — |
| model.safetensors.index.json | Configuration | 189.1 KB | — |
| preprocessor_config.json | Configuration | 488 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| quantization_config.json | Configuration | 314.8 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| README.md | Documentation | 6.8 KB | — |
| chat_template.jinja | Other | 3.9 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 6.8 GB
Released by ramGPT through its official repository on Hugging Face. Read the license.
Built From
- Derived from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- Quantized from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 6.8 GB |
| 16-bit | 6.8 GB |
| 8-bit | 3.4 GB |
| 4-bit | 1.7 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw
How much GPU memory does MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw need?
About 8.2 GB at 16-bit and 2.1 GB at 4-bit: the weights (3.4B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw commercially?
Yes. MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
What is MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.