SAVRN
Search Contact SAVRN

Open-weight model

deepseek-v4.1-flash-gguf

by Salvatore Sanfilippo antirez/deepseek-v4.1-flash-gguf

deepseek-v4.1-flash-gguf is an open-weight model from Salvatore Sanfilippo, released under MIT License. Its published files total 885.3 GB. It draws 1.5M downloads a month.

Calibrated weights for the Metal implementation in DwarfStar, developed on the ds4.1flash branch. These are not interchangeable with DeepSeek V4 Flash weights or its DSpark drafter. The language GGUFs each include 188.83 GiB of native FP8 Engram tables.

Parameters—
Context—
Weights366.7 GB
Licensemit
AccessOpen weights
Monthly Downloads1.5M

Model Card

By Salvatore Sanfilippo, published under mit, revision dd8a266f7145.

Calibrated weights for the Metal implementation in DwarfStar, developed on the ds4.1flash branch. These are not interchangeable with DeepSeek V4 Flash weights or its DSpark drafter. The language GGUFs each include 188.83 GiB of native FP8 Engram tables. DwarfStar reads the needed rows directly from disk; it never makes the whole table resident. Keep the GGUF on a fast local SSD. The main-weight sizes above do not include context or runtime buffers. From a DwarfStar checkout with V4.1 support: Use SSD streaming on a single 128 GB Mac. Q2 also fits resident tensor parallelism across two 128 GB Macs with RDMA, or full residency on a larger Mac. Q4 needs SSD streaming on smaller Macs; the…

Read Salvatore Sanfilippo's full model card

DeepSeek V4.1 Flash GGUF for DwarfStar

Calibrated weights for the Metal implementation in DwarfStar, developed on the ds4.1flash branch. These are not interchangeable with DeepSeek V4 Flash weights or its DSpark drafter.

Download target File size Main weights
ds41f-q2 340.60 GiB 151.77 GiB
ds41f-q4 482.98 GiB 294.15 GiB
ds41f-vision 0.90 GiB Separate vision encoder

The language GGUFs each include 188.83 GiB of native FP8 Engram tables. DwarfStar reads the needed rows directly from disk; it never makes the whole table resident. Keep the GGUF on a fast local SSD. The main-weight sizes above do not include context or runtime buffers.

Download and Run

From a DwarfStar checkout with V4.1 support:

./download_model.sh ds41f-q2
./ds4 -m gguf/DeepSeek-V4.1-Flash-Q2.gguf --ssd-streaming --ctx 32768

Use SSD streaming on a single 128 GB Mac. Q2 also fits resident tensor parallelism across two 128 GB Macs with RDMA, or full residency on a larger Mac. Q4 needs SSD streaming on smaller Macs; the resident target is a 512 GB Mac. It does not fit resident TP across two 128 GB Macs.

./download_model.sh ds41f-q4
./download_model.sh ds41f-vision

Q4 exceeds Hugging Face's single-file limit, so it is stored in two binary parts. The downloader joins them into DeepSeek-V4.1-Flash-Q4.gguf, verifies the complete file and removes the temporary parts. Allow another 37 GiB of free disk space while joining. Rerun the command after an interruption. The individual parts are not runnable GGUFs.

For images, add --vision gguf/DeepSeek-V4.1-Flash-Vision.gguf to the language model command. The same model and memory options work with ds4-agent and ds4-server.

Quantization

Q2 uses IQ2_XXS routed gate/up experts and Q2_K down experts. Q4 uses Q4_K for all routed experts. Both retain Q8 attention projections, shared experts and output, with F16/F32 tensors elsewhere and unchanged native Engram data. They use the same 8,192-token activation imatrix, not a requantization of one GGUF into the other.

Source: deepseek-ai/DeepSeek-V4.1-Flash at df42c109f1defefcbfcedbe7d905718a12266e40. The separate encoder comes from the same checkpoint. See the included upstream MIT license.

Identity and Version

Repository
antirez/deepseek-v4.1-flash-gguf
Publisher
Salvatore Sanfilippo
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
dd8a266f7145edc19e2334b46e19b6821f221dc7
First published
2026-09-11
Last updated
2026-09-12

Files and Weights

7 files, 885.3 GB in total. The weights are 2 files totalling 366.7 GB in gguf.

Weights2 files · 366.7 GB
Documentation2 files · 3.5 KB
Other2 files · 518.6 GB
Repository1 file · 1.8 KB
Every file
FileTypeSizeSHA-256
DeepSeek-V4.1-Flash-Q2.ggufWeights365.7 GB 1ce6a8f88062
DeepSeek-V4.1-Flash-Vision.ggufWeights970.6 MB cc283f032b3e
LICENSEDocumentation1.1 KB —
README.mdDocumentation2.4 KB —
DeepSeek-V4.1-Flash-Q4.gguf.part1Other480.0 GB 6442b1f92240
DeepSeek-V4.1-Flash-Q4.gguf.part2Other38.6 GB 7c3e10646c91
.gitattributesRepository1.8 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
366.7 GB
Download from Salvatore Sanfilippo

Released by Salvatore Sanfilippo through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published366.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About deepseek-v4.1-flash-gguf

Can I use deepseek-v4.1-flash-gguf commercially?

Yes. deepseek-v4.1-flash-gguf is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.