SAVRN
Search Contact SAVRN

Open-weight model

MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw

by ramGPT ramgpt/MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw

MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw is an open-weight model from ramGPT, released under MIT License. It has 3.4B parameters and a 262,144-token context. At 16-bit it needs about 8.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. Validated locally on an NVIDIA GeForce RTX 4090. The throughput figures are short local smoke measurements, not standardized cross-system benchmarks.

Parameters3.4B
Context262,144
Weights6.8 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw (3.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 6.8 GB 8.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 3.4 GB 4.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.7 GB 2.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 6, 2026.

MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw on every accelerator the SAVRN Index prices, at every precision

Model Card

By ramGPT, published under mit, revision 81328546105e.

EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. Validated locally on an NVIDIA GeForce RTX 4090. The throughput figures are short local smoke measurements, not standardized cross-system benchmarks. Tool-call validation produced a structured getweather call for Toronto. Vision validation used a generated test image containing a large red square; the model correctly returned red. The source weights were not modified. Two local metadata normalizations were required for the current ExLlamaV3 conversion path: 1. The Qwen3.5 processor metadata declared Qwen2VLImageProcessor; the local conversion copy was normalized to Qwen2VLImageProcessorFast to match the ExLlamaV3 Qwen3.5…

Read ramGPT's full model card

EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B.

Build provenance

  • Source revision: 2367e865d009c13ac81713a2878291d33ab28177
  • ExLlamaV3: 1.5.2+cu128.torch2.10.0
  • PyTorch: 2.10.0+cu128
  • Target bitrate: 4.00 bpw
  • Architecture: Qwen3_5ForConditionalGeneration
  • Source license: MIT
  • Artifact size: approximately 6.4 GiB

Validation

Validated locally on an NVIDIA GeForce RTX 4090.

Check Result
Direct ExLlamaV3 load + generation PASS
Direct decode smoke 118.009 tok/s
TabbyAPI model load PASS
OpenAI-compatible chat completion PASS
TabbyAPI chat decode smoke 147.64 tok/s
Streaming PASS
Qwen3.5 tool-call parsing PASS
Vision PASS
Two simultaneous API requests PASS
TabbyAPI configured context 8192 tokens

The throughput figures are short local smoke measurements, not standardized cross-system benchmarks.

Tool-call validation produced a structured get_weather call for Toronto.

Vision validation used a generated test image containing a large red square; the model correctly returned red.

Source metadata compatibility notes

The source weights were not modified. Two local metadata normalizations were required for the current ExLlamaV3 conversion path:

  1. The Qwen3.5 processor metadata declared Qwen2VLImageProcessor; the local conversion copy was normalized to Qwen2VLImageProcessorFast to match the ExLlamaV3 Qwen3.5 loader expectation.
  2. The source config declared text_config.mtp_num_hidden_layers = 1, while the source safetensors index contained no MTP/NextN tensors. The local conversion copy therefore disabled the absent MTP side model.

These changes affect conversion metadata only and are recorded in the local build provenance.

TabbyAPI configuration used

model:
  backend: exllamav3
  max_seq_len: 8192
  cache_size: 8192
  gpu_split_auto: false
  gpu_split: [23, 0]
  vision: true
  reasoning: true
  reasoning_start_token: "<think>"
  reasoning_end_token: "</think>"
  tool_format: qwen3_5

The validation environment intentionally excluded the RTX 3060 from model placement so the reported 4090 measurements are not distorted by heterogeneous multi-GPU splitting.

Notes

This is an independent community quantization of the XiaomiMiMo upstream model. Refer to the upstream model card for intended use, limitations, and license terms.

RTX 4090 context scaling

TabbyAPI, RTX 4090 only, EXL3 4.0 bpw, FP16 cache, vision enabled, max batch size 1. Each context test used a unique long prompt containing a hidden needle near the end; the response had to begin with the exact needle code.

Target Actual prompt Prefill tok/s Decode tok/s Total time Needle Observed GPU0 memory
1K 1,038 4,943 145.2 0.91s PASS 8039 MiB
4K 4,104 8,922 144.1 1.09s PASS 8175 MiB
8K 8,199 9,212 140.5 1.45s PASS 8191 MiB
16K 16,392 8,957 137.2 2.37s PASS 8199 MiB
32K 32,541 4,828 127.2 7.31s PASS 8201 MiB
64K 64,587 6,577 111.3 10.48s PASS 8203 MiB

All six needle checks passed through about 64K prompt tokens. Decode throughput fell gradually from about 145 tok/s at 1K to 111 tok/s at 64K.

These are single-pass local smoke measurements, not standardized benchmark scores. Prefill throughput can move non-monotonically with kernel compilation, autotuning, cache state and prompt shape. The GPU-memory column is total observed GPU0 usage during the run, not an isolated model allocation.

Raw results: benchmarks/rtx4090-tabbyapi-context.json.

Long-generation / repetition stability

I also ran six long-generation samples through TabbyAPI:

  • 3 runs with thinking disabled.
  • 3 runs with thinking enabled and a 512-token reasoning budget.
  • Up to 1,024 generated tokens per run.
  • Observed decode throughput stayed around 145-146 tok/s.
  • 0/6 runs triggered the repeated n-gram loop heuristic.
  • The largest repeated 12-word-window count was 2.
  • Repeated 4-gram ratios ranged from 0.0079 to 0.0381.
  • 4/6 runs reached the 1,024-token ceiling, so this should not be interpreted as an EOS-behavior test.

Some outputs repeated short Markdown-formatting lines; those are retained in the raw data but are not counted as semantic repetition loops unless longer n-grams repeat as well.

This is a small stability smoke test, not proof that repetition loops cannot occur with other prompts or sampling settings.

Raw results: benchmarks/rtx4090-loop-stability.json.

Hardware comparison: RTX 4090 vs RTX 3060 vs mixed split

The same EXL3 artifact was tested three ways. These are local TabbyAPI smoke measurements, not standardized benchmark scores.

Placement ~1K decode tok/s ~8K decode tok/s ~16K decode tok/s Short generation tok/s
RTX 4090 only 145.2 140.5 137.3 ~145-148
RTX 4090 + RTX 3060 layer split 106.2 102.4 100.7 108.1
RTX 3060 12GB only 53.5 49.0 50.3 55.8

RTX 3060 12GB only

The model ran entirely on the RTX 3060 with no CPU model offload and no model layers placed on the RTX 4090. Vision was disabled for this accessibility test.

  • 1K prompt: 53.5 tok/s, needle PASS
  • 8K prompt: 49.0 tok/s, needle PASS
  • ~16K prompt: 50.3 tok/s, needle PASS
  • Short normal generation: 55.8 tok/s
  • Repeated 4-gram ratio: 0.0
  • Max repeated 12-gram count: 1
  • Obvious repetition loop detected: no

Observed RTX 3060 memory was about 6.4-6.6 GiB during these runs. That is total observed GPU memory, not a model-only allocation measurement.

RTX 4090 + RTX 3060 mixed layer split

The mixed test used a conventional layer split with gpu_split [4, 8], not tensor parallelism.

  • 1K prompt: 106.2 tok/s
  • 8K prompt: 102.4 tok/s
  • ~16K prompt: 100.7 tok/s
  • Short normal generation: 108.1 tok/s
  • Needle retrieval passed at all three context sizes.
  • No obvious repetition loop was detected in the short generation smoke.

After load, total observed GPU memory was approximately 4.74 GiB on the RTX 4090 and 2.65 GiB on the RTX 3060.

For this 9B model, the mixed layer split is slower than a 4090-only placement because the model already fits comfortably on the faster card. The mixed result is still useful as a proof that heterogeneous layer splitting works and remains usable when a future model no longer fits on one GPU.

Raw results:

  • benchmarks/rtx3060-tabbyapi.json
  • benchmarks/rtx4090-rtx3060-mixed.json

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
4,096
Feed-forward size
12,288
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5
Quantization
exl3

Identity and Version

Repository
ramgpt/MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw
Publisher
ramGPT
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
3.4B parameters
Languages
Not stated by the source
Revision
81328546105e5906af880578aecade33e2292383
First published
2026-09-30
Last updated
2026-09-30

Files and Weights

20 files, 6.9 GB in total. The weights are 2 files totalling 6.8 GB in safetensors.

Weights2 files · 6.8 GB
Configuration11 files · 521.1 KB
Tokenizer4 files · 30.1 MB
Documentation1 file · 6.8 KB
Other1 file · 3.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights4.2 GB 8262ad181237
model-00002-of-00002.safetensorsWeights2.6 GB 819f9002d461
benchmarks/rtx3060-tabbyapi.jsonConfiguration2.9 KB —
benchmarks/rtx4090-loop-stability.jsonConfiguration2.9 KB —
benchmarks/rtx4090-rtx3060-mixed.jsonConfiguration2.7 KB —
benchmarks/rtx4090-tabbyapi-context.jsonConfiguration2.9 KB —
config.jsonConfiguration3.6 KB —
generation_config.jsonConfiguration126 B —
model.safetensors.index.jsonConfiguration189.1 KB —
preprocessor_config.jsonConfiguration488 B —
processor_config.jsonConfiguration1.2 KB —
quantization_config.jsonConfiguration314.8 KB —
video_preprocessor_config.jsonConfiguration385 B —
README.mdDocumentation6.8 KB —
chat_template.jinjaOther3.9 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
mit
Access
Open weights, no gate
Download size
6.8 GB
Download from ramGPT

Released by ramGPT through its official repository on Hugging Face. Read the license.

Built From

  • Derived from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
  • Quantized from XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

Memory Requirements

PrecisionWeights in memory
As published6.8 GB
16-bit6.8 GB
8-bit3.4 GB
4-bit1.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw

How much GPU memory does MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw need?

About 8.2 GB at 16-bit and 2.1 GB at 4-bit: the weights (3.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw commercially?

Yes. MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is MiMo-V2.6-Distill-Qwen-9B-EXL3-4.0bpw's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.