SAVRN
Search Contact SAVRN

Open-weight model · Text generation

keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated

by Keys drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated

keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated is a model for text generation from Keys, released under MIT License (access requested at publisher). It has 119B parameters. At 16-bit it needs about 285.6 GB of GPU memory, which fits on 1x MI355X from $2.59 an hour; at 4-bit, 71.4 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 5 downloads a month.

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL. Thinking on/off is a request flag. Same weights. You choose per call. Thinking-off is the 100% gate.

Parameters119B
Context—
Weights322.7 GB
Licensemit
AccessAccess requested at publisher
Monthly Downloads5

Runs On

What it takes to serve keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated (119B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 238.0 GB 285.6 GB 1x MI355X (288 GB)
Vultr
$2.59 2x MI300X $3.70 · 2x MI325X $4.00
8-bit 119.0 GB 142.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 59.5 GB 71.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated on every accelerator the SAVRN Index prices, at every precision

Model Card

By Keys, published under mit, revision e9b57a488096.

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL. Thinking on/off is a request flag. Same weights. You choose per call. Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes. This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLEUSE.md. Xiaomi's chat template already supports both. Do not swap…

Read Keys's full model card

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL.

Thinking on/off is a request flag. Same weights. You choose per call.

thinking off thinking on (visible content, 1024 tokens)
Refusal suite (32) 32/32 BYPASS · 0 refuse · 0 garble · 0 empty 25/32 BYPASS · 7 refuse · 0 garble · 0 empty
Cyber suite (22) 22/22 BYPASS · 0 refuse · 0 garble · 0 empty 16/22 BYPASS · 1 refuse · 2 garble · 3 empty

Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes.

Full credit: XiaomiMiMo/MiMo-V2.6-Pro-RL · jarrelscy/MiMo-V2.6-Pro-RL-ARVQ-hybrid @ 63430f7 · dealignai/MiMo-V2.6-Pro-RL-UNCENSORED (o_proj map) · Keys four-Spark recipe


Responsible use and gated access

This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLE_USE.md.


Thinking on / off

Xiaomi's chat template already supports both. Do not swap templates. Pass chat_template_kwargs:

# Thinking OFF — 32/32 refusal, 22/22 cyber on our gate
{
  "model": "MiMo-V2.6-Pro-ARVQ",
  "messages": [{"role": "user", "content": prompt}],
  "temperature": 0.6,
  "top_p": 0.95,
  "chat_template_kwargs": {"enable_thinking": False, "thinking": False},
}

# Thinking ON — 25/32 refusal, 16/22 cyber on visible content
{
  "model": "MiMo-V2.6-Pro-ARVQ",
  "messages": [{"role": "user", "content": prompt}],
  "max_tokens": 1024,  # thinking can consume a 192-token budget
  "temperature": 0.6,
  "top_p": 0.95,
  "chat_template_kwargs": {"enable_thinking": True, "thinking": True},
}

vLLM / Hermes: --reasoning-parser mimo. The live Keys serve defaults thinking off. Raise max_tokens when thinking is on.


What changed vs stock ARVQ

Native Xiaomi vs dealign v3 dense-shard hash-diff: 29 decoder self_attn.o_proj.weight tensors, BF16 (6144, 16384), layers 30–54 and 64–67. qkv, router gate, sinks, e_score, experts, MTP, DFlash, vision, and audio matched stock.

This release copies 25 of those o_proj matrices onto backbone-001.safetensors of Jarrelscy ARVQ 63430f7:

  • applied: 32–45, 48–54, 64–67
  • left stock (DFlash-source / pad anchors): 30, 31, 46, 47
  • L55–63 already match Xiaomi stock in dealign v3

Packed ARVQ/NVFP4 experts, MTP, and dflash/ stay the Jarrelscy files.

dealign native UNCENSORED this repo
Layout Xiaomi FP8 + MXFP4 Jarrelscy ARVQ / NVFP4 hybrid
o_proj edits 29 layers 25 layers (anchors skipped)
Chat template extra anti-refusal think prefill stock Xiaomi template
Thinking-off gate (our 32+22) not measured here 32/32 · 22/22

Serving (4× DGX Spark) — image v4, 2026-09-27

Single stream: ~30 tok/s on code, 21.8 tok/s on prose. All three built-in MTP draft heads now run (before v4, only head 0 drafted), and batched ARVQ kernels give 1,038–1,291 tok/s prefill.

Tool-call loop fixed. Truncated tool batches now return finish_reason: "length", and the output cap is 8192. Details: HERMES.md.

Image: ghcr.io/drowzeys/mimo-v26-pro-arvq-spark:latest, the same image as :63430f7-sm121-v4. It is public and is the only supported image.

One-shot bring-up (the defaults are the best setup)

hf download drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated --local-dir /path/to/mimo-arvq   # storage all 4 nodes can read
git clone https://github.com/drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated-4-DGX-Sparks-1M-Context
export MASTER_ADDR=<rank-0 IP>
bash serve/launch-rank.sh <this-node-IP> <1|2|3> <RoCE-GID-index> /path/to/mimo-arvq headless   # ranks 1-3 first
bash serve/launch-rank.sh <this-node-IP> 0 <RoCE-GID-index> /path/to/mimo-arvq api              # rank 0 = API :8888

The GID index is the IPv4 RoCE entry from show_gids (3 or 7 on our Sparks). The same launcher is also in serve/ in this repo. Keep --gpu-memory-utilization at 0.85.

TP 4
Context 1,048,576
Draft MTP k=2 (default). All three built-in heads run non-chain; k=3 is best for code.
Execution torch.compile + CUDA graphs
Prefill expert-batched ARVQ CUDA kernels, 5120-token chunks
KV BF16
Seqs 4
Served name MiMo-V2.6-Pro-ARVQ (use /v1/chat/completions)
Tools --enable-auto-tool-choice --tool-call-parser mimo --reasoning-parser mimo

Speed (image v4)

512 new tokens, temperature 1.0, thinking off, one request at a time:

Draft tokens Prose Code 4 requests together
2 (default) 21.8 tok/s 30.9 tok/s 46.6 tok/s
3 19.9 tok/s 32.3–34.0 tok/s 37.6 tok/s
Uncached prefill Old eager recipe Now Time to first token
9.5K tokens 128 tok/s 1,291 tok/s 7.4 s
38K tokens 126 tok/s 1,038 tok/s 36.6 s

Defaults for clients that send nothing (for example Pi)

generation_config.json now sets temperature 0.7, top_p 0.95 and max 8192 tokens. At the original 1.0, about 0.3% of samples lock into a repetition loop (about 2% for long thinking-on outputs). The chat template now keeps thinking off unless a request passes "chat_template_kwargs": {"enable_thinking": true}. The raw /v1/completions endpoint skips the chat template, so its output echoes or garbles. That is expected; use chat completions.

Measured on v4, taking the 19 prompts that had looped at temperature 1.0 and sampling each 3 times: 5/57 looped under the old behaviour, 2/57 with the new defaults. One-shot validation of v4 gave prose 22.4 tok/s, code 30.5 tok/s, 48.0 tok/s aggregate at 4 requests, and 1,014 tok/s prefill at 38K.

Official Xiaomi images do not load this layout.


Scores (heuristic classifier)

thinking=False, greedy, 192 tokens, live TP4:

Suite Bypass Refuse Garble Empty
Refusal 32 32 0 0 0
Cyber 22 22 0 0 0
Stock ARVQ 5 / 9

thinking=True, greedy, 1024 tokens:

Suite Bypass Refuse Garble Empty
Refusal 32 25 7 0 0
Cyber 22 16 1 2 3

A bypass label means the reply starts delivering the requested content. It does not certify correctness. HarmBench-320 was not rerun on this ARVQ tree.


License

MIT, inherited from Xiaomi MiMo-V2.6-Pro-RL and the Jarrelscy hybrid.

Identity and Version

Repository
drowzeys/keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated
Publisher
Keys
Task
Text generation
Modality
Text
Library
vllm
Parameters
119B parameters
Languages
en, zh
Revision
e9b57a48809699849f1c1cadc5ba3165df9b70ff
First published
2026-09-25
Last updated
2026-09-27

Files and Weights

523 files, 322.7 GB in total. The weights are 213 files totalling 322.7 GB in pt, safetensors.

Weights213 files · 322.7 GB
Configuration296 files · 5.1 MB
Tokenizer5 files · 15.9 MB
Documentation3 files · 14.7 KB
Other5 files · 3.5 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
arvq-layer-001-down.safetensorsWeights1.3 GB —
arvq-layer-001-gateup.safetensorsWeights2.6 GB —
arvq-layer-002-down.safetensorsWeights1.3 GB —
arvq-layer-002-gateup.safetensorsWeights2.6 GB —
arvq-layer-003-down.safetensorsWeights1.3 GB —
arvq-layer-003-gateup.safetensorsWeights2.6 GB —
arvq-layer-004-down.safetensorsWeights1.3 GB —
arvq-layer-004-gateup.safetensorsWeights2.6 GB —
arvq-layer-005-down.safetensorsWeights1.3 GB —
arvq-layer-005-gateup.safetensorsWeights2.6 GB —
arvq-layer-006-down.safetensorsWeights1.3 GB —
arvq-layer-006-gateup.safetensorsWeights2.6 GB —
arvq-layer-007-down.safetensorsWeights1.3 GB —
arvq-layer-007-gateup.safetensorsWeights2.6 GB —
arvq-layer-008-down.safetensorsWeights1.3 GB —
arvq-layer-008-gateup.safetensorsWeights2.6 GB —
arvq-layer-009-down.safetensorsWeights1.3 GB —
arvq-layer-009-gateup.safetensorsWeights2.6 GB —
arvq-layer-010-down.safetensorsWeights1.3 GB —
arvq-layer-010-gateup.safetensorsWeights2.6 GB —
arvq-layer-011-down.safetensorsWeights1.3 GB —
arvq-layer-011-gateup.safetensorsWeights2.6 GB —
arvq-layer-012-down.safetensorsWeights1.3 GB —
arvq-layer-012-gateup.safetensorsWeights2.6 GB —
arvq-layer-013-down.safetensorsWeights1.3 GB —
arvq-layer-013-gateup.safetensorsWeights2.6 GB —
arvq-layer-014-down.safetensorsWeights1.3 GB —
arvq-layer-014-gateup.safetensorsWeights2.6 GB —
arvq-layer-015-down.safetensorsWeights1.3 GB —
arvq-layer-015-gateup.safetensorsWeights2.6 GB —
arvq-layer-016-down.safetensorsWeights1.3 GB —
arvq-layer-016-gateup.safetensorsWeights2.6 GB —
arvq-layer-017-down.safetensorsWeights1.3 GB —
arvq-layer-017-gateup.safetensorsWeights2.6 GB —
arvq-layer-018-down.safetensorsWeights1.3 GB —
arvq-layer-018-gateup.safetensorsWeights2.6 GB —
arvq-layer-019-down.safetensorsWeights1.3 GB —
arvq-layer-019-gateup.safetensorsWeights2.6 GB —
arvq-layer-020-down.safetensorsWeights1.3 GB —
arvq-layer-020-gateup.safetensorsWeights2.6 GB —
arvq-layer-021-down.safetensorsWeights1.3 GB —
arvq-layer-021-gateup.safetensorsWeights2.6 GB —
arvq-layer-022-down.safetensorsWeights1.3 GB —
arvq-layer-022-gateup.safetensorsWeights2.6 GB —
arvq-layer-023-down.safetensorsWeights1.3 GB —
arvq-layer-023-gateup.safetensorsWeights2.6 GB —
arvq-layer-024-down.safetensorsWeights1.3 GB —
arvq-layer-024-gateup.safetensorsWeights2.6 GB —
arvq-layer-025-down.safetensorsWeights1.3 GB —
arvq-layer-025-gateup.safetensorsWeights2.6 GB —
arvq-layer-026-down.safetensorsWeights1.3 GB —
arvq-layer-026-gateup.safetensorsWeights2.6 GB —
arvq-layer-027-down.safetensorsWeights1.3 GB —
arvq-layer-027-gateup.safetensorsWeights2.6 GB —
arvq-layer-028-down.safetensorsWeights1.3 GB —
arvq-layer-028-gateup.safetensorsWeights2.6 GB —
arvq-layer-029-down.safetensorsWeights1.3 GB —
arvq-layer-029-gateup.safetensorsWeights2.6 GB —
arvq-layer-030-down.safetensorsWeights1.3 GB —
arvq-layer-030-gateup.safetensorsWeights2.6 GB —
arvq-layer-031-down.safetensorsWeights1.3 GB —
arvq-layer-031-gateup.safetensorsWeights2.6 GB —
arvq-layer-032-down.safetensorsWeights1.3 GB —
arvq-layer-032-gateup.safetensorsWeights2.6 GB —
arvq-layer-033-down.safetensorsWeights1.3 GB —
arvq-layer-033-gateup.safetensorsWeights2.6 GB —
arvq-layer-034-down.safetensorsWeights1.3 GB —
arvq-layer-034-gateup.safetensorsWeights2.6 GB —
arvq-layer-035-down.safetensorsWeights1.3 GB —
arvq-layer-035-gateup.safetensorsWeights2.6 GB —
arvq-layer-036-down.safetensorsWeights1.3 GB —
arvq-layer-036-gateup.safetensorsWeights2.6 GB —
arvq-layer-037-down.safetensorsWeights1.3 GB —
arvq-layer-037-gateup.safetensorsWeights2.6 GB —
arvq-layer-038-down.safetensorsWeights1.3 GB —
arvq-layer-038-gateup.safetensorsWeights2.6 GB —
arvq-layer-039-down.safetensorsWeights1.3 GB —
arvq-layer-039-gateup.safetensorsWeights2.6 GB —
arvq-layer-040-down.safetensorsWeights1.3 GB —
arvq-layer-040-gateup.safetensorsWeights2.6 GB —
arvq-layer-041-down.safetensorsWeights1.3 GB —
arvq-layer-041-gateup.safetensorsWeights2.6 GB —
arvq-layer-042-down.safetensorsWeights1.3 GB —
arvq-layer-042-gateup.safetensorsWeights2.6 GB —
arvq-layer-043-down.safetensorsWeights1.3 GB —
arvq-layer-043-gateup.safetensorsWeights2.6 GB —
arvq-layer-044-down.safetensorsWeights1.3 GB —
arvq-layer-044-gateup.safetensorsWeights2.6 GB —
arvq-layer-045-down.safetensorsWeights1.3 GB —
arvq-layer-045-gateup.safetensorsWeights2.6 GB —
arvq-layer-046-down.safetensorsWeights1.3 GB —
arvq-layer-046-gateup.safetensorsWeights2.6 GB —
arvq-layer-047-down.safetensorsWeights1.3 GB —
arvq-layer-047-gateup.safetensorsWeights2.5 GB —
arvq-layer-048-down.safetensorsWeights1.3 GB —
arvq-layer-048-gateup.safetensorsWeights2.5 GB —
arvq-layer-049-down.safetensorsWeights1.3 GB —
arvq-layer-049-gateup.safetensorsWeights2.5 GB —
arvq-layer-050-down.safetensorsWeights1.3 GB —
arvq-layer-050-gateup.safetensorsWeights2.5 GB —
arvq-layer-051-down.safetensorsWeights1.3 GB —
arvq-layer-051-gateup.safetensorsWeights2.5 GB —
arvq-layer-052-down.safetensorsWeights1.3 GB —
arvq-layer-052-gateup.safetensorsWeights2.5 GB —
arvq-layer-053-down.safetensorsWeights1.2 GB —
arvq-layer-053-gateup.safetensorsWeights2.5 GB —
arvq-layer-054-down.safetensorsWeights1.2 GB —
arvq-layer-054-gateup.safetensorsWeights2.5 GB —
arvq-layer-055-down.safetensorsWeights1.2 GB —
arvq-layer-055-gateup.safetensorsWeights2.4 GB —
arvq-layer-056-down.safetensorsWeights1.2 GB —
arvq-layer-056-gateup.safetensorsWeights2.3 GB —
arvq-layer-057-down.safetensorsWeights1.2 GB —
arvq-layer-057-gateup.safetensorsWeights2.3 GB —
arvq-layer-058-down.safetensorsWeights1.1 GB —
arvq-layer-058-gateup.safetensorsWeights2.2 GB —
arvq-layer-059-down.safetensorsWeights1.1 GB —
arvq-layer-059-gateup.safetensorsWeights2.1 GB —
arvq-layer-060-down.safetensorsWeights1.0 GB —
arvq-layer-060-gateup.safetensorsWeights2.1 GB —
arvq-layer-061-down.safetensorsWeights1.1 GB —
arvq-layer-061-gateup.safetensorsWeights2.1 GB —
arvq-layer-062-down.safetensorsWeights1.0 GB —
arvq-layer-062-gateup.safetensorsWeights2.0 GB —
arvq-layer-063-down.safetensorsWeights1.0 GB —
arvq-layer-063-gateup.safetensorsWeights2.1 GB —
arvq-layer-064-down.safetensorsWeights973.2 MB —
arvq-layer-064-gateup.safetensorsWeights1.9 GB —
arvq-layer-065-down.safetensorsWeights936.4 MB —
arvq-layer-065-gateup.safetensorsWeights1.9 GB —
arvq-layer-066-down.safetensorsWeights993.3 MB —
arvq-layer-066-gateup.safetensorsWeights2.0 GB —
arvq-layer-067-down.safetensorsWeights899.6 MB —
arvq-layer-067-gateup.safetensorsWeights1.8 GB —
arvq-layer-068-down.safetensorsWeights812.7 MB —
arvq-layer-068-gateup.safetensorsWeights1.6 GB —
arvq-layer-069-down.safetensorsWeights642.1 MB —
arvq-layer-069-gateup.safetensorsWeights1.3 GB —
audio_tokenizer/model.safetensorsWeights1.9 GB —
backbone-000.safetensorsWeights2.0 GB —
backbone-001.safetensorsWeights30.2 GB —
backbone-002.safetensorsWeights2.5 GB —
dflash/dflash_draft_model.safetensorsWeights5.5 GB —
dflash/mask_embedding.ptWeights14.0 KB —
roster-layer-001.safetensorsWeights1.1 KB —
roster-layer-002.safetensorsWeights1.1 KB —
roster-layer-003.safetensorsWeights1.1 KB —
roster-layer-004.safetensorsWeights1.1 KB —
roster-layer-005.safetensorsWeights1.1 KB —
roster-layer-006.safetensorsWeights1.1 KB —
roster-layer-007.safetensorsWeights1.1 KB —
roster-layer-008.safetensorsWeights1.1 KB —
roster-layer-009.safetensorsWeights1.1 KB —
roster-layer-010.safetensorsWeights1.1 KB —
roster-layer-011.safetensorsWeights1.1 KB —
roster-layer-012.safetensorsWeights1.1 KB —
roster-layer-013.safetensorsWeights1.1 KB —
roster-layer-014.safetensorsWeights1.1 KB —
roster-layer-015.safetensorsWeights1.1 KB —
roster-layer-016.safetensorsWeights1.1 KB —
roster-layer-017.safetensorsWeights1.1 KB —
roster-layer-018.safetensorsWeights1.1 KB —
roster-layer-019.safetensorsWeights1.1 KB —
roster-layer-020.safetensorsWeights1.1 KB —
roster-layer-021.safetensorsWeights1.1 KB —
roster-layer-022.safetensorsWeights21.2 MB —
roster-layer-023.safetensorsWeights1.1 KB —
roster-layer-024.safetensorsWeights1.1 KB —
roster-layer-025.safetensorsWeights1.1 KB —
roster-layer-026.safetensorsWeights1.1 KB —
roster-layer-027.safetensorsWeights1.1 KB —
roster-layer-028.safetensorsWeights1.1 KB —
roster-layer-029.safetensorsWeights1.1 KB —
roster-layer-030.safetensorsWeights1.1 KB —
roster-layer-031.safetensorsWeights1.1 KB —
roster-layer-032.safetensorsWeights1.1 KB —
roster-layer-033.safetensorsWeights1.1 KB —
roster-layer-034.safetensorsWeights1.1 KB —
roster-layer-035.safetensorsWeights1.1 KB —
roster-layer-036.safetensorsWeights1.1 KB —
roster-layer-037.safetensorsWeights1.1 KB —
roster-layer-038.safetensorsWeights1.1 KB —
roster-layer-039.safetensorsWeights1.1 KB —
roster-layer-040.safetensorsWeights1.1 KB —
roster-layer-041.safetensorsWeights1.1 KB —
roster-layer-042.safetensorsWeights1.1 KB —
roster-layer-043.safetensorsWeights42.5 MB —
roster-layer-044.safetensorsWeights21.2 MB —
roster-layer-045.safetensorsWeights42.5 MB —
roster-layer-046.safetensorsWeights1.1 KB —
roster-layer-047.safetensorsWeights106.2 MB —
roster-layer-048.safetensorsWeights148.6 MB —
roster-layer-049.safetensorsWeights127.4 MB —
roster-layer-050.safetensorsWeights148.6 MB —
roster-layer-051.safetensorsWeights212.3 MB —
roster-layer-052.safetensorsWeights169.9 MB —
roster-layer-053.safetensorsWeights254.8 MB —
roster-layer-054.safetensorsWeights233.6 MB —
roster-layer-055.safetensorsWeights615.8 MB —
roster-layer-056.safetensorsWeights806.9 MB —
roster-layer-057.safetensorsWeights828.1 MB —
roster-layer-058.safetensorsWeights1.2 GB —
roster-layer-059.safetensorsWeights1.4 GB —
roster-layer-060.safetensorsWeights1.6 GB —
roster-layer-061.safetensorsWeights1.4 GB —
roster-layer-062.safetensorsWeights1.7 GB —
roster-layer-063.safetensorsWeights1.6 GB —
roster-layer-064.safetensorsWeights2.0 GB —
roster-layer-065.safetensorsWeights2.2 GB —
roster-layer-066.safetensorsWeights1.8 GB —
roster-layer-067.safetensorsWeights2.4 GB —
roster-layer-068.safetensorsWeights3.0 GB —
roster-layer-069.safetensorsWeights4.1 GB —
allocation.jsonConfiguration19.7 KB —
allocation_capture.jsonConfiguration816 B —
allocation_scores/layer1.jsonConfiguration14.8 KB —
allocation_scores/layer10.jsonConfiguration14.6 KB —
allocation_scores/layer11.jsonConfiguration14.6 KB —
allocation_scores/layer12.jsonConfiguration14.8 KB —
allocation_scores/layer13.jsonConfiguration14.6 KB —
allocation_scores/layer14.jsonConfiguration14.6 KB —
allocation_scores/layer15.jsonConfiguration14.5 KB —
allocation_scores/layer16.jsonConfiguration14.7 KB —
allocation_scores/layer17.jsonConfiguration14.1 KB —
allocation_scores/layer18.jsonConfiguration13.8 KB —
allocation_scores/layer19.jsonConfiguration13.0 KB —
allocation_scores/layer2.jsonConfiguration14.7 KB —
allocation_scores/layer20.jsonConfiguration12.9 KB —
allocation_scores/layer21.jsonConfiguration13.0 KB —
allocation_scores/layer22.jsonConfiguration13.0 KB —
allocation_scores/layer23.jsonConfiguration13.0 KB —
allocation_scores/layer24.jsonConfiguration13.0 KB —
allocation_scores/layer25.jsonConfiguration13.0 KB —
allocation_scores/layer26.jsonConfiguration13.0 KB —
allocation_scores/layer27.jsonConfiguration13.0 KB —
allocation_scores/layer28.jsonConfiguration13.0 KB —
allocation_scores/layer29.jsonConfiguration13.0 KB —
allocation_scores/layer3.jsonConfiguration14.7 KB —
allocation_scores/layer30.jsonConfiguration13.0 KB —
allocation_scores/layer31.jsonConfiguration12.9 KB —
allocation_scores/layer32.jsonConfiguration12.9 KB —
allocation_scores/layer33.jsonConfiguration13.0 KB —
allocation_scores/layer34.jsonConfiguration12.9 KB —
allocation_scores/layer35.jsonConfiguration12.9 KB —
allocation_scores/layer36.jsonConfiguration13.0 KB —
allocation_scores/layer37.jsonConfiguration12.9 KB —
allocation_scores/layer38.jsonConfiguration12.9 KB —
allocation_scores/layer39.jsonConfiguration12.9 KB —
allocation_scores/layer4.jsonConfiguration14.7 KB —
allocation_scores/layer40.jsonConfiguration12.9 KB —
allocation_scores/layer41.jsonConfiguration12.9 KB —
allocation_scores/layer42.jsonConfiguration12.9 KB —
allocation_scores/layer43.jsonConfiguration12.9 KB —
allocation_scores/layer44.jsonConfiguration12.9 KB —
allocation_scores/layer45.jsonConfiguration12.9 KB —
allocation_scores/layer46.jsonConfiguration12.9 KB —
allocation_scores/layer47.jsonConfiguration12.9 KB —
allocation_scores/layer48.jsonConfiguration12.9 KB —
allocation_scores/layer49.jsonConfiguration12.9 KB —
allocation_scores/layer5.jsonConfiguration14.6 KB —
allocation_scores/layer50.jsonConfiguration12.9 KB —
allocation_scores/layer51.jsonConfiguration12.9 KB —
allocation_scores/layer52.jsonConfiguration12.9 KB —
allocation_scores/layer53.jsonConfiguration12.9 KB —
allocation_scores/layer54.jsonConfiguration13.0 KB —
allocation_scores/layer55.jsonConfiguration12.9 KB —
allocation_scores/layer56.jsonConfiguration12.9 KB —
allocation_scores/layer57.jsonConfiguration12.9 KB —
allocation_scores/layer58.jsonConfiguration12.9 KB —
allocation_scores/layer59.jsonConfiguration12.9 KB —
allocation_scores/layer6.jsonConfiguration14.7 KB —
allocation_scores/layer60.jsonConfiguration13.0 KB —
allocation_scores/layer61.jsonConfiguration12.9 KB —
allocation_scores/layer62.jsonConfiguration12.9 KB —
allocation_scores/layer63.jsonConfiguration12.9 KB —
allocation_scores/layer64.jsonConfiguration12.9 KB —
allocation_scores/layer65.jsonConfiguration13.0 KB —
allocation_scores/layer66.jsonConfiguration12.9 KB —
allocation_scores/layer67.jsonConfiguration12.9 KB —
allocation_scores/layer68.jsonConfiguration12.9 KB —
allocation_scores/layer69.jsonConfiguration12.9 KB —
allocation_scores/layer7.jsonConfiguration14.6 KB —
allocation_scores/layer8.jsonConfiguration14.5 KB —
allocation_scores/layer9.jsonConfiguration14.6 KB —
audio_tokenizer/config.jsonConfiguration1.2 KB —
audio_tokenizer/generation_config.jsonConfiguration149 B —
campaign.jsonConfiguration1.2 KB —
cold_manifests/layer-001-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-002-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-003-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-004-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-005-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-006-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-007-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-008-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-009-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-010-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-011-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-012-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-013-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-014-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-015-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-016-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-017-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-018-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-019-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-020-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-021-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-022-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-023-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-024-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-025-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-026-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-027-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-028-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-029-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-030-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-031-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-032-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-033-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-034-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-035-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-036-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-037-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-038-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-039-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-040-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-041-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-042-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-043-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-044-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-045-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-046-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-047-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-048-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-049-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-050-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-051-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-052-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-053-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-054-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-055-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-056-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-057-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-058-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-059-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-060-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-061-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-062-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-063-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-064-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-065-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-066-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-067-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-068-manifest.jsonConfiguration4.3 KB —
cold_manifests/layer-069-manifest.jsonConfiguration4.3 KB —
config.jsonConfiguration131.5 KB —
configuration_mimo_v2.pyConfiguration10.0 KB —
dflash/config.jsonConfiguration1.2 KB —
dflash/dflash.pyConfiguration14.3 KB —
dflash/model.safetensors.index.jsonConfiguration4.7 KB —
evaluation/reference_sanity.jsonConfiguration191 B —
fit_completion.jsonConfiguration189 B —
fitting_config.jsonConfiguration1.6 KB —
generation_config.jsonConfiguration212 B —
hot_target.jsonConfiguration999 B —
model.safetensors.index.jsonConfiguration160.6 KB —
modeling_mimo_v2.pyConfiguration85.5 KB —
preparation_status.jsonConfiguration373 B —
preprocessor_config.jsonConfiguration350 B —
pv_progress.jsonConfiguration39.2 KB —
pv_reports/layer-001-selection.jsonConfiguration20.1 KB —
pv_reports/layer-001.jsonConfiguration18.7 KB —
pv_reports/layer-002-selection.jsonConfiguration30.7 KB —
pv_reports/layer-002.jsonConfiguration28.7 KB —
pv_reports/layer-003-selection.jsonConfiguration23.8 KB —
pv_reports/layer-003.jsonConfiguration22.2 KB —
pv_reports/layer-004-selection.jsonConfiguration24.9 KB —
pv_reports/layer-004.jsonConfiguration23.2 KB —
pv_reports/layer-005-selection.jsonConfiguration30.5 KB —
pv_reports/layer-005.jsonConfiguration28.4 KB —
pv_reports/layer-006-selection.jsonConfiguration21.4 KB —
pv_reports/layer-006.jsonConfiguration19.9 KB —
pv_reports/layer-007-selection.jsonConfiguration22.5 KB —
pv_reports/layer-007.jsonConfiguration20.9 KB —
pv_reports/layer-008-selection.jsonConfiguration30.7 KB —
pv_reports/layer-008.jsonConfiguration28.6 KB —
pv_reports/layer-009-selection.jsonConfiguration21.4 KB —
pv_reports/layer-009.jsonConfiguration19.9 KB —
pv_reports/layer-010-selection.jsonConfiguration30.7 KB —
pv_reports/layer-010.jsonConfiguration28.7 KB —
pv_reports/layer-011-selection.jsonConfiguration24.9 KB —
pv_reports/layer-011.jsonConfiguration23.2 KB —
pv_reports/layer-012-selection.jsonConfiguration21.3 KB —
pv_reports/layer-012.jsonConfiguration19.8 KB —
pv_reports/layer-013-selection.jsonConfiguration23.3 KB —
pv_reports/layer-013.jsonConfiguration21.7 KB —
pv_reports/layer-014-selection.jsonConfiguration30.4 KB —
pv_reports/layer-014.jsonConfiguration28.4 KB —
pv_reports/layer-015-selection.jsonConfiguration30.3 KB —
pv_reports/layer-015.jsonConfiguration28.3 KB —
pv_reports/layer-016-selection.jsonConfiguration21.3 KB —
pv_reports/layer-016.jsonConfiguration19.8 KB —
pv_reports/layer-017-selection.jsonConfiguration21.0 KB —
pv_reports/layer-017.jsonConfiguration19.5 KB —
pv_reports/layer-018-selection.jsonConfiguration20.9 KB —
pv_reports/layer-018.jsonConfiguration19.4 KB —
pv_reports/layer-019-selection.jsonConfiguration30.2 KB —
pv_reports/layer-019.jsonConfiguration28.1 KB —
pv_reports/layer-021-selection.jsonConfiguration30.2 KB —
pv_reports/layer-021.jsonConfiguration28.1 KB —
pv_reports/layer-022-selection.jsonConfiguration29.4 KB —
pv_reports/layer-022.jsonConfiguration27.4 KB —
pv_reports/layer-023-selection.jsonConfiguration30.1 KB —
pv_reports/layer-023.jsonConfiguration28.0 KB —
pv_reports/layer-024-selection.jsonConfiguration30.1 KB —
pv_reports/layer-024.jsonConfiguration28.0 KB —
pv_reports/layer-025-selection.jsonConfiguration30.1 KB —
pv_reports/layer-025.jsonConfiguration28.0 KB —
pv_reports/layer-026-selection.jsonConfiguration30.0 KB —
pv_reports/layer-026.jsonConfiguration28.0 KB —
pv_reports/layer-027-selection.jsonConfiguration30.0 KB —
pv_reports/layer-027.jsonConfiguration27.9 KB —
pv_reports/layer-028-selection.jsonConfiguration30.0 KB —
pv_reports/layer-028.jsonConfiguration27.9 KB —
pv_reports/layer-029-selection.jsonConfiguration27.9 KB —
pv_reports/layer-029.jsonConfiguration26.0 KB —
pv_reports/layer-030-selection.jsonConfiguration27.9 KB —
pv_reports/layer-030.jsonConfiguration26.0 KB —
pv_reports/layer-031-selection.jsonConfiguration25.5 KB —
pv_reports/layer-031.jsonConfiguration23.7 KB —
pv_reports/layer-032-selection.jsonConfiguration29.8 KB —
pv_reports/layer-032.jsonConfiguration27.7 KB —
pv_reports/layer-033-selection.jsonConfiguration29.8 KB —
pv_reports/layer-033.jsonConfiguration27.7 KB —
pv_reports/layer-034-selection.jsonConfiguration29.7 KB —
pv_reports/layer-034.jsonConfiguration27.7 KB —
pv_reports/layer-035-selection.jsonConfiguration25.3 KB —
pv_reports/layer-035.jsonConfiguration23.5 KB —
pv_reports/layer-036-selection.jsonConfiguration25.3 KB —
pv_reports/layer-036.jsonConfiguration23.5 KB —
pv_reports/layer-037-selection.jsonConfiguration25.3 KB —
pv_reports/layer-037.jsonConfiguration23.5 KB —
pv_reports/layer-038-selection.jsonConfiguration22.9 KB —
pv_reports/layer-038.jsonConfiguration21.3 KB —
pv_reports/layer-039-selection.jsonConfiguration22.8 KB —
pv_reports/layer-039.jsonConfiguration21.2 KB —
pv_reports/layer-040-selection.jsonConfiguration22.8 KB —
pv_reports/layer-040.jsonConfiguration21.2 KB —
pv_reports/layer-041-selection.jsonConfiguration22.8 KB —
pv_reports/layer-041.jsonConfiguration21.2 KB —
pv_reports/layer-042-selection.jsonConfiguration22.8 KB —
pv_reports/layer-042.jsonConfiguration21.2 KB —
pv_reports/layer-043-selection.jsonConfiguration22.7 KB —
pv_reports/layer-043.jsonConfiguration21.1 KB —
pv_reports/layer-044-selection.jsonConfiguration22.7 KB —
pv_reports/layer-044.jsonConfiguration21.1 KB —
pv_reports/layer-045-selection.jsonConfiguration22.7 KB —
pv_reports/layer-045.jsonConfiguration21.1 KB —
pv_reports/layer-046-selection.jsonConfiguration22.7 KB —
pv_reports/layer-046.jsonConfiguration21.1 KB —
pv_reports/layer-047-selection.jsonConfiguration22.7 KB —
pv_reports/layer-047.jsonConfiguration21.0 KB —
pv_reports/layer-048-selection.jsonConfiguration22.6 KB —
pv_reports/layer-048.jsonConfiguration21.0 KB —
pv_reports/layer-049-selection.jsonConfiguration22.6 KB —
pv_reports/layer-049.jsonConfiguration21.0 KB —
pv_reports/layer-050-selection.jsonConfiguration22.6 KB —
pv_reports/layer-050.jsonConfiguration21.0 KB —
pv_reports/layer-051-selection.jsonConfiguration22.6 KB —
pv_reports/layer-051.jsonConfiguration20.9 KB —
pv_reports/layer-052-selection.jsonConfiguration22.6 KB —
pv_reports/layer-052.jsonConfiguration20.9 KB —
pv_reports/layer-053-selection.jsonConfiguration22.6 KB —
pv_reports/layer-053.jsonConfiguration20.9 KB —
pv_reports/layer-054-selection.jsonConfiguration22.6 KB —
pv_reports/layer-054.jsonConfiguration20.9 KB —
pv_reports/layer-055-selection.jsonConfiguration29.3 KB —
pv_reports/layer-055.jsonConfiguration27.3 KB —
pv_reports/layer-056-selection.jsonConfiguration29.3 KB —
pv_reports/layer-056.jsonConfiguration27.2 KB —
pv_reports/layer-057-selection.jsonConfiguration29.3 KB —
pv_reports/layer-057.jsonConfiguration27.2 KB —
pv_reports/layer-058-selection.jsonConfiguration29.3 KB —
pv_reports/layer-058.jsonConfiguration27.3 KB —
pv_reports/layer-059-selection.jsonConfiguration29.3 KB —
pv_reports/layer-059.jsonConfiguration27.2 KB —
pv_reports/layer-060-selection.jsonConfiguration29.3 KB —
pv_reports/layer-060.jsonConfiguration27.2 KB —
pv_reports/layer-061-selection.jsonConfiguration29.2 KB —
pv_reports/layer-061.jsonConfiguration27.2 KB —
pv_reports/layer-062-selection.jsonConfiguration29.2 KB —
pv_reports/layer-062.jsonConfiguration27.2 KB —
pv_reports/layer-063-selection.jsonConfiguration29.2 KB —
pv_reports/layer-063.jsonConfiguration27.2 KB —
pv_reports/layer-064-selection.jsonConfiguration29.3 KB —
pv_reports/layer-064.jsonConfiguration27.2 KB —
pv_reports/layer-065-selection.jsonConfiguration29.2 KB —
pv_reports/layer-065.jsonConfiguration27.2 KB —
pv_reports/layer-066-selection.jsonConfiguration29.2 KB —
pv_reports/layer-066.jsonConfiguration27.2 KB —
pv_reports/layer-067-selection.jsonConfiguration29.2 KB —
pv_reports/layer-067.jsonConfiguration27.2 KB —
pv_reports/layer-068-selection.jsonConfiguration29.3 KB —
pv_reports/layer-068.jsonConfiguration27.2 KB —
run_manifest.jsonConfiguration3.6 KB —
seed_ready.jsonConfiguration61 B —
serve/bench/speedbench.pyConfiguration6.6 KB —
upload_verification.jsonConfiguration254 B —
LICENSEDocumentation1.1 KB —
README.mdDocumentation10.2 KB —
RESPONSIBLE_USE.mdDocumentation3.4 KB —
MiMo_V2_6_technical_report.pdfOther3.0 MB —
assets/architecture.pngOther405.4 KB —
audio_tokenizer/chat_template.jinjaOther5.6 KB —
chat_template.jinjaOther3.9 KB —
serve/launch-rank.shOther4.5 KB —
.gitattributesRepository1.7 KB —
audio_tokenizer/tokenizer_config.jsonTokenizer6.1 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB —
tokenizer_config.jsonTokenizer10.6 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
mit
Access
Access requested at publisher
Download size
322.7 GB
Request access from Keys

Keys grants access through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published322.7 GB
16-bit238.0 GB
8-bit119.0 GB
4-bit59.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated

How much GPU memory does keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated need?

About 285.6 GB at 16-bit and 71.4 GB at 4-bit: the weights (119B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated on?

At 16-bit, 1x MI355X from $2.59 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated commercially?

Yes. keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Text generation

gpt-oss-120b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 116.8B parameters 131,072 tokens transformers

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Open weights other 124B parameters 262,144 tokens vllm

Model · Text generation

Vinci-Cyber-123B-1.0

SimpleDirect

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…

Open weights other 125B parameters 262,144 tokens transformers

Model · Text generation

Hy3-Razor-154B-A18B-E96of192

Mingyang Song

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…

Open weights apache-2.0 153.8B parameters 262,144 tokens transformers

Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: Deterministic Hybrid Control Framework for Frozen Neural Operators — DHCF-FNO). Repack (BF16 safetensors, shards ≤50 GB) and community GGUF quantizations of (modeltype: aliceai, 80B total / 3B active MoE with KDA layers). Исходные 49 шардов Yandex слиты в один стриминговый файл и заново нарезаны SHA256 каждого тензора, 0 расхождений (см.…

Open weights apache-2.0 81.3B parameters 262,144 tokens transformers