SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

DeepSeek-V4.1-Flash-lossless-CSF

by Local Inference Lab local-inference-lab/DeepSeek-V4.1-Flash-lossless-CSF

DeepSeek-V4.1-Flash-lossless-CSF is an open-weight model for image and text to text from Local Inference Lab, released under MIT License. Its published files total 495.7 GB.

A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4.1-Flash at revision dba1be0a40aa45a94ad051997016db3960a90277. CSF stores the routed-expert block-scale planes in a compressed form.

Parameters—
Context—
Weights495.6 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By Local Inference Lab, published under mit, revision c5c41fe301c4.

A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4.1-Flash at revision dba1be0a40aa45a94ad051997016db3960a90277. CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly. - Built with trellis-quant trellisquant.losslessscalecheckpoint (commit 60feca330087, family deepseekv41). - verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files…

Read Local Inference Lab's full model card

A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4.1-Flash at revision dba1be0a40aa45a94ad051997016db3960a90277.

CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly.

Routed experts (40 layers x 384) - MXFP4 (E2M1, UE8M0 scale per 32)
Engram tables, attention, shared experts - FP8 E4M3 with UE8M0 scales
Vision encoder, aligner, MTP (including its routed experts) - source format

Routed-expert block scales - lossless MXFP4-CSF (row-base-offset1-u24-exceptions/1)

Sizes

  • Weight files: 510.30 GB in the source, 495.63 GB here (14.67 GB saved).
  • Compressed scales: 46,080 matrices (UE8M0 block scales of the 40 x 384 x 3 main-layer routed-expert projections): 16.99 GB -> 2.31 GB (13.6%).

Provenance

  • Source: deepseek-ai/DeepSeek-V4.1-Flash revision dba1be0a40aa45a94ad051997016db3960a90277, 48 index-referenced safetensors shards. During export, each source shard's SHA-256 was checked against its Hugging Face LFS hash.
  • Built with trellis-quant trellis_quant.lossless_scale_checkpoint (commit 60feca330087, family deepseek_v41).
  • verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files matched their SHA-256 (46,080 scale matrices, 48 shards, passed).
  • Hub main (2cba9e42) adds chat_template.jinja and is otherwise byte-identical to dba1be0a. Every tensor and config file matches.
  • tensors/ is byte-identical (all 48 compressed-shard SHA-256 values) to local-inference-lab/DeepSeek-V4.1-Flash-MXFP4-CSF revision 872da235. This build adds the source LICENSE to metadata/ and records the source revision.

Layout

lil-mxfp4-csf-checkpoint/1, codec row-base-offset1-u24-exceptions/1:

  • tensors/ - the source shard names; each routed-expert scale <name> is stored as <name>.mxfp4_csf_fixed (uint8) plus <name>.mxfp4_csf_exceptions (uint32)
  • metadata/ - byte copies of the source's config, tokenizer, index, README and LICENSE
  • manifest.json, build-contract.json, receipts/ (per-shard source headers and hashes), verification.json, LICENSE

There is no top-level model.safetensors.index.json, so a plain safetensors loader will not open this directory by mistake.

Serving

Use vLLM with the MXFP4-CSF reader: --quantization mxfp4_csf --load-format mxfp4_csf. Point vLLM at a serving directory that holds the files of metadata/, with config.json's quantization_config replaced by:

{
  "...": "every key of metadata/config.json quantization_config",
  "quant_method": "mxfp4_csf",
  "format_version": 1,
  "checkpoint_root": "/path/to/this/checkpoint"
}

The weights are read from checkpoint_root; the serving directory holds only metadata.

Verify or restore

PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint verify \
  --checkpoint /path/to/this/checkpoint --workers 16
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint restore \
  --checkpoint /path/to/this/checkpoint --destination /path/to/fresh/dir --workers 16

verify rebuilds every original shard and compares SHA-256 values. restore writes the original shards and metadata files into a fresh directory, checking every hash before a file is published. The result is an ordinary copy of the source checkpoint.

License

Same license as the source; LICENSE is copied unchanged from deepseek-ai/DeepSeek-V4.1-Flash.

Identity and Version

Repository
local-inference-lab/DeepSeek-V4.1-Flash-lossless-CSF
Publisher
Local Inference Lab
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
csf
Revision
c5c41fe301c4c1d24fc09376883d9168e521dc66
First published
2026-10-04
Last updated
2026-10-04

Files and Weights

108 files, 495.7 GB in total. The weights are 48 files totalling 495.6 GB in safetensors.

Weights48 files · 495.6 GB
Configuration53 files · 43.4 MB
Tokenizer2 files · 6.4 MB
Documentation4 files · 19.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
tensors/model-00001-of-00048.safetensorsWeights970.5 MB 7a5035831774
tensors/model-00002-of-00048.safetensorsWeights1.3 GB 7582cad4247e
tensors/model-00003-of-00048.safetensorsWeights7.0 GB ce24f95e0345
tensors/model-00004-of-00048.safetensorsWeights7.0 GB e897f4f9106e
tensors/model-00005-of-00048.safetensorsWeights7.0 GB 25e4b6929a16
tensors/model-00006-of-00048.safetensorsWeights7.0 GB 7fcb54a73b4d
tensors/model-00007-of-00048.safetensorsWeights7.0 GB c54f51df209f
tensors/model-00008-of-00048.safetensorsWeights7.0 GB e379683c8c33
tensors/model-00009-of-00048.safetensorsWeights7.0 GB 26f274864a45
tensors/model-00010-of-00048.safetensorsWeights7.0 GB 54a7b1b0423f
tensors/model-00011-of-00048.safetensorsWeights7.0 GB a68790a34752
tensors/model-00012-of-00048.safetensorsWeights7.0 GB 8102c986824e
tensors/model-00013-of-00048.safetensorsWeights7.0 GB e16918d2bb60
tensors/model-00014-of-00048.safetensorsWeights7.0 GB 59adfabb37db
tensors/model-00015-of-00048.safetensorsWeights7.0 GB d47caadae0a9
tensors/model-00016-of-00048.safetensorsWeights7.0 GB 912813c38687
tensors/model-00017-of-00048.safetensorsWeights7.0 GB 614a771280cc
tensors/model-00018-of-00048.safetensorsWeights7.0 GB a4418f6ce676
tensors/model-00019-of-00048.safetensorsWeights7.0 GB 484b00e8ce26
tensors/model-00020-of-00048.safetensorsWeights7.0 GB 9b3293beb180
tensors/model-00021-of-00048.safetensorsWeights7.0 GB 5b5cd9d2fa0d
tensors/model-00022-of-00048.safetensorsWeights7.0 GB a8ae9f775d24
tensors/model-00023-of-00048.safetensorsWeights7.0 GB 26043b3d4e98
tensors/model-00024-of-00048.safetensorsWeights7.0 GB 18d77548abf6
tensors/model-00025-of-00048.safetensorsWeights7.0 GB b9a081ebf2c0
tensors/model-00026-of-00048.safetensorsWeights7.0 GB 06ba91a8b97c
tensors/model-00027-of-00048.safetensorsWeights7.0 GB 78a9bb6bb3d2
tensors/model-00028-of-00048.safetensorsWeights7.0 GB a033ca2a78f4
tensors/model-00029-of-00048.safetensorsWeights7.0 GB 512e788ee155
tensors/model-00030-of-00048.safetensorsWeights7.0 GB 1f21961b7c84
tensors/model-00031-of-00048.safetensorsWeights7.0 GB ab2e549e5edb
tensors/model-00032-of-00048.safetensorsWeights7.0 GB fb19798058a3
tensors/model-00033-of-00048.safetensorsWeights7.0 GB 5d04e3e43078
tensors/model-00034-of-00048.safetensorsWeights7.0 GB de07bb1bf5fd
tensors/model-00035-of-00048.safetensorsWeights7.0 GB 5451b7142e72
tensors/model-00036-of-00048.safetensorsWeights7.0 GB a4c906f8ff00
tensors/model-00037-of-00048.safetensorsWeights7.0 GB dca721bce557
tensors/model-00038-of-00048.safetensorsWeights7.0 GB ad2a325bf903
tensors/model-00039-of-00048.safetensorsWeights7.0 GB d134f54f5a7b
tensors/model-00040-of-00048.safetensorsWeights7.0 GB d3691d6ffa4b
tensors/model-00041-of-00048.safetensorsWeights7.0 GB ae31755dfdd7
tensors/model-00042-of-00048.safetensorsWeights7.0 GB 8bc460ca51e5
tensors/model-00043-of-00048.safetensorsWeights1.3 GB bed0ea6a4c3e
tensors/model-00044-of-00048.safetensorsWeights2.7 GB 41ba4af7d57d
tensors/model-00045-of-00048.safetensorsWeights2.6 GB 1ee4a0103fad
tensors/model-00046-of-00048.safetensorsWeights2.7 GB 7e5333799713
tensors/model-00047-of-00048.safetensorsWeights101.5 GB 18523c0ef7ad
tensors/model-00048-of-00048.safetensorsWeights101.5 GB d5d45a5fadd4
build-contract.jsonConfiguration7.5 MB —
manifest.jsonConfiguration16.9 KB —
metadata/config.jsonConfiguration3.3 KB —
metadata/model.safetensors.index.jsonConfiguration7.5 MB —
receipts/model-00001-of-00048.safetensors.jsonConfiguration36.9 KB —
receipts/model-00002-of-00048.safetensors.jsonConfiguration854 B —
receipts/model-00003-of-00048.safetensors.jsonConfiguration696.8 KB —
receipts/model-00004-of-00048.safetensors.jsonConfiguration696.5 KB —
receipts/model-00005-of-00048.safetensors.jsonConfiguration697.7 KB —
receipts/model-00006-of-00048.safetensors.jsonConfiguration696.4 KB —
receipts/model-00007-of-00048.safetensors.jsonConfiguration696.4 KB —
receipts/model-00008-of-00048.safetensors.jsonConfiguration696.4 KB —
receipts/model-00009-of-00048.safetensors.jsonConfiguration696.4 KB —
receipts/model-00010-of-00048.safetensors.jsonConfiguration696.4 KB —
receipts/model-00011-of-00048.safetensors.jsonConfiguration697.7 KB —
receipts/model-00012-of-00048.safetensors.jsonConfiguration696.4 KB —
receipts/model-00013-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00014-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00015-of-00048.safetensors.jsonConfiguration700.8 KB —
receipts/model-00016-of-00048.safetensors.jsonConfiguration700.9 KB —
receipts/model-00017-of-00048.safetensors.jsonConfiguration702.2 KB —
receipts/model-00018-of-00048.safetensors.jsonConfiguration700.8 KB —
receipts/model-00019-of-00048.safetensors.jsonConfiguration700.9 KB —
receipts/model-00020-of-00048.safetensors.jsonConfiguration701.0 KB —
receipts/model-00021-of-00048.safetensors.jsonConfiguration701.0 KB —
receipts/model-00022-of-00048.safetensors.jsonConfiguration700.9 KB —
receipts/model-00023-of-00048.safetensors.jsonConfiguration701.9 KB —
receipts/model-00024-of-00048.safetensors.jsonConfiguration700.8 KB —
receipts/model-00025-of-00048.safetensors.jsonConfiguration700.9 KB —
receipts/model-00026-of-00048.safetensors.jsonConfiguration700.8 KB —
receipts/model-00027-of-00048.safetensors.jsonConfiguration701.3 KB —
receipts/model-00028-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00029-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00030-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00031-of-00048.safetensors.jsonConfiguration701.2 KB —
receipts/model-00032-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00033-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00034-of-00048.safetensors.jsonConfiguration700.8 KB —
receipts/model-00035-of-00048.safetensors.jsonConfiguration701.2 KB —
receipts/model-00036-of-00048.safetensors.jsonConfiguration700.7 KB —
receipts/model-00037-of-00048.safetensors.jsonConfiguration700.8 KB —
receipts/model-00038-of-00048.safetensors.jsonConfiguration701.0 KB —
receipts/model-00039-of-00048.safetensors.jsonConfiguration701.3 KB —
receipts/model-00040-of-00048.safetensors.jsonConfiguration700.9 KB —
receipts/model-00041-of-00048.safetensors.jsonConfiguration701.2 KB —
receipts/model-00042-of-00048.safetensors.jsonConfiguration701.6 KB —
receipts/model-00043-of-00048.safetensors.jsonConfiguration630 B —
receipts/model-00044-of-00048.safetensors.jsonConfiguration113.6 KB —
receipts/model-00045-of-00048.safetensors.jsonConfiguration113.2 KB —
receipts/model-00046-of-00048.safetensors.jsonConfiguration114.5 KB —
receipts/model-00047-of-00048.safetensors.jsonConfiguration1.3 KB —
receipts/model-00048-of-00048.safetensors.jsonConfiguration1.3 KB —
verification.jsonConfiguration14.9 KB —
LICENSEDocumentation1.1 KB —
README.mdDocumentation4.0 KB —
metadata/LICENSEDocumentation1.1 KB —
metadata/README.mdDocumentation13.1 KB —
.gitattributesRepository1.5 KB —
metadata/tokenizer.jsonTokenizer6.4 MB —
metadata/tokenizer_config.jsonTokenizer801 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
495.6 GB
Download from Local Inference Lab

Released by Local Inference Lab through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published495.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DeepSeek-V4.1-Flash-lossless-CSF

Can I use DeepSeek-V4.1-Flash-lossless-CSF commercially?

Yes. DeepSeek-V4.1-Flash-lossless-CSF is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…

Open weights apache-2.0 gguf

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. BF16/16-bit MTP GGUF also avail. (there is also a extra robust "tools" version too - Q6 and Q8.) Extreme intelligence in a small package. Jaw dropping performance. Superior…

Open weights apache-2.0