SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

DeepSeek-V4.1-MLX-Q9

by Inferencer inferencerlabs/DeepSeek-V4.1-MLX-Q9

DeepSeek-V4.1-MLX-Q9 is an open-weight model for image and text to text from Inferencer. It has 1,048,576-token context. Its published files total 1.4 MB.

M3 Ultra 512GB + M4 Max 128GB: ~15 tokens/s @ 1000 tokens ~501 GiB This build repacks the majority of the model's existing quantization, with the remainder quantized to 9 bits. This configuration empirically yielded the highest quality in our testing.

Parameters—
Context1,048,576
Weights1.4 MB
License—
AccessOpen weights
Monthly Downloads—

Model Card

M3 Ultra 512GB + M4 Max 128GB: ~15 tokens/s @ 1000 tokens ~501 GiB This build repacks the majority of the model's existing quantization, with the remainder quantized to 9 bits. This configuration empirically yielded the highest quality in our testing. We are not the creator, originator, or owner of any model listed. Each model is created and provided by third parties. Models may not always be accurate or contextually appropriate. You are responsible for verifying the information before making important decisions. We are not liable for any damages, losses, or issues arising from its use, including data loss or inaccuracies in AI-generated content.

Excerpt from the card by Inferencer.

Configuration

Architecture
DeepseekV41ForCausalLM
Context length (tokens)
1,048,576
Layers
40
Hidden size
5,120
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
384
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Model type
deepseek_v41

Identity and Version

Repository
inferencerlabs/DeepSeek-V4.1-MLX-Q9
Publisher
Inferencer
Task
Image and text to text
Modality
Image and text
Library
mlx
Parameters
Not stated by the source
Languages
en
Revision
a0de49f925c0d023c23329906aad061588c0424d
First published
2026-10-04
Last updated
2026-10-04

Files and Weights

38 files, 1.4 MB in total.

Configuration17 files · 575.1 KB
Tokenizer1 file · 801 B
Documentation5 files · 19.7 KB
Other14 files · 773.1 KB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
config.jsonConfiguration389.9 KB —
encoding/encoding.pyConfiguration35.3 KB —
encoding/test_encoding.pyConfiguration14.3 KB —
encoding/tests/test_input_1.jsonConfiguration2.8 KB —
encoding/tests/test_input_2.jsonConfiguration527 B —
encoding/tests/test_input_3.jsonConfiguration2.6 KB —
encoding/tests/test_input_4.jsonConfiguration712 B —
encoding/tests/test_input_5.jsonConfiguration1.1 KB —
inference/config.jsonConfiguration2.0 KB —
inference/convert.pyConfiguration9.5 KB —
inference/engram.pyConfiguration8.1 KB —
inference/examples/example_harmony.jsonConfiguration2.2 KB —
inference/generate.pyConfiguration8.7 KB —
inference/image_processor.pyConfiguration7.7 KB —
inference/kernel.pyConfiguration23.8 KB —
inference/model.pyConfiguration61.5 KB —
inference/vision.pyConfiguration4.5 KB —
LICENSEDocumentation1.1 KB —
README.mdDocumentation1.7 KB —
encoding/README.mdDocumentation11.0 KB —
evaluation/README.mdDocumentation4.0 KB —
inference/README.mdDocumentation2.0 KB —
assets/dsv41_agentic_performance.pngOther190.7 KB 44deae01cb9c
assets/dsv41_kv_cache.pngOther270.9 KB b61bf4651d4b
chat_template.jinjaOther5.7 KB —
encoding/tests/test_output_1.txtOther2.5 KB —
encoding/tests/test_output_2.txtOther294 B —
encoding/tests/test_output_3.txtOther2.5 KB —
encoding/tests/test_output_4.txtOther574 B —
encoding/tests/test_output_5.txtOther408 B —
evaluation/dsh-minimal.patchOther28.7 KB —
inference/examples/example.txtOther332 B —
inference/examples/images/carrots.jpegOther212.5 KB 5df896a4a07e
inference/examples/images/corn.jpegOther56.1 KB —
inference/requirements.txtOther97 B —
inference/run.shOther1.8 KB —
.gitattributesRepository1.7 KB —
tokenizer_config.jsonTokenizer801 B —

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download from Inferencer

Released by Inferencer through its official repository on Hugging Face.

Built From

Questions About DeepSeek-V4.1-MLX-Q9

What is DeepSeek-V4.1-MLX-Q9's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…

Open weights apache-2.0 gguf

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. BF16/16-bit MTP GGUF also avail. (there is also a extra robust "tools" version too - Q6 and Q8.) Extreme intelligence in a small package. Jaw dropping performance. Superior…

Open weights apache-2.0