with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence. Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base. Every response 4-tier graded (HARDREF / SOFTRED / HEDGE /…
Open-weight model · Image and text to text
DeepSeek-V4.1-Flash
by DeepSeek deepseek-ai/DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash is an open-weight model for image and text to text from DeepSeek, released under MIT License. It has 763.2B parameters and a 1,048,576-token context. At 16-bit it needs about 1831.7 GB of GPU memory, which fits on 8x MI325X from $16.00 an hour; at 4-bit, 457.9 GB on 2x MI325X from $4.00, at the lowest prices in the SAVRN Index. It draws 1.2M downloads a month.
We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture.
Runs On
What it takes to serve DeepSeek-V4.1-Flash (763.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 1526.4 GB | 1831.7 GB | 8x MI325X (256 GB) Vultr |
$16.00 | 7x MI355X $18.13 · 7x B300 $46.20 |
| 8-bit | 763.2 GB | 915.8 GB | 4x MI325X (256 GB) Vultr |
$8.00 | 5x MI300X $9.25 · 4x MI355X $10.36 |
| 4-bit | 381.6 GB | 457.9 GB | 2x MI325X (256 GB) Vultr |
$4.00 | 2x MI355X $5.18 · 3x MI300X $5.55 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
DeepSeek-V4.1-Flash on every accelerator the SAVRN Index prices, at every precision
Model Card
By DeepSeek, published under mit, revision 2cba9e42aa02.
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
Introduction
We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.
Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads. SWA Bounded Replay reconstructs missing SWA KV states by replaying only the most recent n_win tokens, avoiding the need to persist SWA KV to SSD and reducing the persistent KV cache footprint to roughly 1/8 of that of DeepSeek-V4-Flash.
Configuration
- Architecture
- DeepseekV41ForCausalLM
- Context length (tokens)
- 1,048,576
- Layers
- 40
- Hidden size
- 5,120
- Attention heads
- 64
- Key/value heads
- 1
- Head dimension
- 512
- Vocabulary size
- 129,280
- Routed experts
- 384
- Experts active per token
- 6
- Sliding window (tokens)
- 128
- RoPE base
- 10,000
- Model type
- deepseek_v41
- Quantization
- fp8
Identity and Version
- Repository
- deepseek-ai/DeepSeek-V4.1-Flash
- Publisher
- DeepSeek
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 763.2B parameters
- Languages
- Not stated by the source
- Revision
- 2cba9e42aa026125f3ed06c6d98c1db82f7ca027
- First published
- 2026-09-10
- Last updated
- 2026-10-01
Files and Weights
89 files, 510.3 GB in total. The weights are 48 files totalling 510.3 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00048.safetensors | Weights | 970.5 MB | 886aebdafa08 |
| model-00002-of-00048.safetensors | Weights | 1.3 GB | 4320066fc695 |
| model-00003-of-00048.safetensors | Weights | 7.4 GB | e1281f85d0ce |
| model-00004-of-00048.safetensors | Weights | 7.4 GB | 79456c9db0cd |
| model-00005-of-00048.safetensors | Weights | 7.4 GB | 4a42dc78698b |
| model-00006-of-00048.safetensors | Weights | 7.4 GB | 020a6df51a28 |
| model-00007-of-00048.safetensors | Weights | 7.4 GB | 40f8b52f763f |
| model-00008-of-00048.safetensors | Weights | 7.4 GB | d62cca4e698f |
| model-00009-of-00048.safetensors | Weights | 7.4 GB | 1ca62e4c294d |
| model-00010-of-00048.safetensors | Weights | 7.4 GB | dd33c9750a40 |
| model-00011-of-00048.safetensors | Weights | 7.4 GB | a9b309f90e0d |
| model-00012-of-00048.safetensors | Weights | 7.4 GB | b359227eceb3 |
| model-00013-of-00048.safetensors | Weights | 7.4 GB | 41d87a4c81fe |
| model-00014-of-00048.safetensors | Weights | 7.4 GB | e7ca4a12688a |
| model-00015-of-00048.safetensors | Weights | 7.4 GB | fa9d49314bbb |
| model-00016-of-00048.safetensors | Weights | 7.4 GB | 07d08bce9d73 |
| model-00017-of-00048.safetensors | Weights | 7.4 GB | 3d35e330a6c2 |
| model-00018-of-00048.safetensors | Weights | 7.4 GB | 48bd0c28b7f4 |
| model-00019-of-00048.safetensors | Weights | 7.4 GB | b7a25cb64a95 |
| model-00020-of-00048.safetensors | Weights | 7.4 GB | 8aea8c4026ba |
| model-00021-of-00048.safetensors | Weights | 7.4 GB | 4cb6558dedfd |
| model-00022-of-00048.safetensors | Weights | 7.4 GB | 81031b68c967 |
| model-00023-of-00048.safetensors | Weights | 7.4 GB | 096723fc8afa |
| model-00024-of-00048.safetensors | Weights | 7.4 GB | c9438ff607bd |
| model-00025-of-00048.safetensors | Weights | 7.4 GB | 4227ef9fe34d |
| model-00026-of-00048.safetensors | Weights | 7.4 GB | 0af8c8f1b96b |
| model-00027-of-00048.safetensors | Weights | 7.4 GB | 3066bd030437 |
| model-00028-of-00048.safetensors | Weights | 7.4 GB | 9bc915075568 |
| model-00029-of-00048.safetensors | Weights | 7.4 GB | d153dd9cde7c |
| model-00030-of-00048.safetensors | Weights | 7.4 GB | 3c9ccd96e908 |
| model-00031-of-00048.safetensors | Weights | 7.4 GB | 0de7b6d7142d |
| model-00032-of-00048.safetensors | Weights | 7.4 GB | 6a5aaa73c6f9 |
| model-00033-of-00048.safetensors | Weights | 7.4 GB | 386e3e91f7f0 |
| model-00034-of-00048.safetensors | Weights | 7.4 GB | 9deab3c4f27c |
| model-00035-of-00048.safetensors | Weights | 7.4 GB | 226573bc07f3 |
| model-00036-of-00048.safetensors | Weights | 7.4 GB | 90d6a85c1eb0 |
| model-00037-of-00048.safetensors | Weights | 7.4 GB | 207aef18f995 |
| model-00038-of-00048.safetensors | Weights | 7.4 GB | cbbaa0b08073 |
| model-00039-of-00048.safetensors | Weights | 7.4 GB | f4cf191547b5 |
| model-00040-of-00048.safetensors | Weights | 7.4 GB | e991bfc41605 |
| model-00041-of-00048.safetensors | Weights | 7.4 GB | 48a1c08afadf |
| model-00042-of-00048.safetensors | Weights | 7.4 GB | e1a4d5d30ae5 |
| model-00043-of-00048.safetensors | Weights | 1.3 GB | d762b688f138 |
| model-00044-of-00048.safetensors | Weights | 2.7 GB | 9a6b39fb88a2 |
| model-00045-of-00048.safetensors | Weights | 2.6 GB | 0cc9d5f6ca3a |
| model-00046-of-00048.safetensors | Weights | 2.7 GB | e625902027b9 |
| model-00047-of-00048.safetensors | Weights | 101.5 GB | 824db4881320 |
| model-00048-of-00048.safetensors | Weights | 101.5 GB | 976330f49543 |
| config.json | Configuration | 3.3 KB | — |
| encoding/encoding.py | Configuration | 37.3 KB | — |
| encoding/test_encoding.py | Configuration | 19.4 KB | — |
| encoding/tests/test_input_1.json | Configuration | 2.8 KB | — |
| encoding/tests/test_input_2.json | Configuration | 527 B | — |
| encoding/tests/test_input_3.json | Configuration | 2.6 KB | — |
| encoding/tests/test_input_4.json | Configuration | 712 B | — |
| encoding/tests/test_input_5.json | Configuration | 1.1 KB | — |
| inference/config.json | Configuration | 2.0 KB | — |
| inference/convert.py | Configuration | 9.5 KB | — |
| inference/engram.py | Configuration | 8.1 KB | — |
| inference/examples/example_harmony.json | Configuration | 2.2 KB | — |
| inference/generate.py | Configuration | 8.7 KB | — |
| inference/image_processor.py | Configuration | 7.7 KB | — |
| inference/kernel.py | Configuration | 23.8 KB | — |
| inference/model.py | Configuration | 61.5 KB | — |
| inference/vision.py | Configuration | 4.5 KB | — |
| model.safetensors.index.json | Configuration | 7.5 MB | — |
| LICENSE | Documentation | 1.1 KB | — |
| README.md | Documentation | 13.1 KB | — |
| encoding/README.md | Documentation | 12.1 KB | — |
| evaluation/README.md | Documentation | 4.0 KB | — |
| inference/README.md | Documentation | 2.0 KB | — |
| DeepSeek_V41_Tech_Report.pdf | Other | 1.8 MB | ba68e2e40408 |
| assets/dsv41_agentic_performance.png | Other | 190.7 KB | 44deae01cb9c |
| assets/dsv41_kv_cache.png | Other | 270.9 KB | b61bf4651d4b |
| chat_template.jinja | Other | 15.3 KB | — |
| encoding/tests/test_output_1.txt | Other | 2.5 KB | — |
| encoding/tests/test_output_2.txt | Other | 294 B | — |
| encoding/tests/test_output_3.txt | Other | 2.5 KB | — |
| encoding/tests/test_output_4.txt | Other | 574 B | — |
| encoding/tests/test_output_5.txt | Other | 408 B | — |
| evaluation/dsh-minimal.patch | Other | 28.7 KB | — |
| inference/examples/example.txt | Other | 332 B | — |
| inference/examples/images/carrots.jpeg | Other | 212.5 KB | 5df896a4a07e |
| inference/examples/images/corn.jpeg | Other | 56.1 KB | — |
| inference/requirements.txt | Other | 97 B | — |
| inference/run.sh | Other | 1.8 KB | — |
| .gitattributes | Repository | 1.7 KB | — |
| tokenizer.json | Tokenizer | 6.4 MB | — |
| tokenizer_config.json | Tokenizer | 801 B | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 510.3 GB
Released by DeepSeek through its official repository on Hugging Face. Read the license.
Evaluations
Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondComparison conditions not established | 90.9 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| cais/hle | Task hleMetric hleSetup With tools; harness not specified in the model card.Comparison conditions not established | 63.9 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| cais/hle | Task hleMetric hleComparison conditions not established | 36.8 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| datacurve/deep-swe | Task deep_sweMetric deep_sweSetup Reported as 'DeepSWE v1.1'; official mini-swe-agent harness.Comparison conditions not established | 74.2 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| harborframework/terminal-bench | Task terminalbench_3Metric terminalbench_3Setup DeepSeek Harness, Minimal mode, 1M-token context.Comparison conditions not established | 30 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| harborframework/terminal-bench | Task terminalbench_4Metric terminalbench_4Setup DeepSeek Harness, Minimal mode, 1M-token context.Comparison conditions not established | 31.2 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| harborframework/terminal-bench-2.1 | Task terminalbench_2_1Metric terminalbench_2_1Setup DeepSeek Harness, Minimal mode, 1M-token context.Comparison conditions not established | 90.6 | DeepSeek-V4.1-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| llamaindex/ExtractBench | Task longMetric longSetup Pipeline name: deepseek_v4_1_flash_extract_oneshot_structured_output_file (served via the DeepSeek API, thinking disabled)Comparison conditions not established | 22.23 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-09-11 |
| llamaindex/ExtractBench | Task meanMetric meanSetup Pipeline name: deepseek_v4_1_flash_extract_oneshot_structured_output_file (served via the DeepSeek API, thinking disabled)Comparison conditions not established | 87.11 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-09-11 |
| llamaindex/ExtractBench | Task mediumMetric mediumSetup Pipeline name: deepseek_v4_1_flash_extract_oneshot_structured_output_file (served via the DeepSeek API, thinking disabled)Comparison conditions not established | 81.49 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-09-11 |
| llamaindex/ExtractBench | Task shortMetric shortSetup Pipeline name: deepseek_v4_1_flash_extract_oneshot_structured_output_file (served via the DeepSeek API, thinking disabled)Comparison conditions not established | 94.44 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-09-11 |
| llamaindex/ParseBench | Task chartMetric chartSetup Pipeline name: deepseek_v4_1_flash_no_thinking_parse_with_layout (served via the DeepSeek API)Comparison conditions not established | 20.46 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| llamaindex/ParseBench | Task layoutMetric layoutSetup Pipeline name: deepseek_v4_1_flash_no_thinking_parse_with_layout (served via the DeepSeek API)Comparison conditions not established | 28.61 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| llamaindex/ParseBench | Task meanMetric meanSetup Pipeline name: deepseek_v4_1_flash_no_thinking_parse_with_layout (served via the DeepSeek API)Comparison conditions not established | 56.57 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| llamaindex/ParseBench | Task tableMetric tableSetup Pipeline name: deepseek_v4_1_flash_no_thinking_parse_with_layout (served via the DeepSeek API)Comparison conditions not established | 79.65 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| llamaindex/ParseBench | Task text_contentMetric text_contentSetup Pipeline name: deepseek_v4_1_flash_no_thinking_parse_with_layout (served via the DeepSeek API)Comparison conditions not established | 88.09 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
| llamaindex/ParseBench | Task text_formattingMetric text_formattingSetup Pipeline name: deepseek_v4_1_flash_no_thinking_parse_with_layout (served via the DeepSeek API)Comparison conditions not established | 66.05 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-09-10 |
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 510.3 GB |
| 16-bit | 1526.4 GB |
| 8-bit | 763.2 GB |
| 4-bit | 381.6 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Hosted Prices
| Host | Input / output | Unit | Observed |
|---|---|---|---|
| Baseten | $0.30 / $1.20 | input / output, per million tokens | Oct 6, 2026 |
| DeepInfra | $0.20 / $0.60 | input / output, per million tokens | Oct 7, 2026 |
| Fireworks | $0.30 / $1.20 | input / output, per million tokens | Oct 6, 2026 |
| Novita | $0.30 / $1.20 | input / output, per million tokens | Oct 7, 2026 |
From the SAVRN Index.
Built on This Model
- Quantized fromdeepseek-v4.1-flash-gguf
- Derived fromdeepseek-v4.1-flash-gguf
- Quantized fromDeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000
- Derived fromDeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000
- Quantized fromzralddeepseekv4.1
- Derived fromzralddeepseekv4.1
- Quantized fromDeepSeek-V4.1-Flash-Abliterated
- Derived fromDeepSeek-V4.1-Flash-Abliterated
- Quantized fromYoungAi-DeepSeek-V4.1-Flash
- Derived fromYoungAi-DeepSeek-V4.1-Flash
- Quantized fromDeepSeek-V4.1-Flash-UNCENSORED-FP8
- Derived fromDeepSeek-V4.1-Flash-UNCENSORED-FP8
- Quantized fromDeepSeek-V4.1-MLX-Q9
- Derived fromDeepSeek-V4.1-MLX-Q9
- Quantized fromDeepSeek-V4.1-Flash-lossless-CSF
- Derived fromDeepSeek-V4.1-Flash-lossless-CSF
Questions About DeepSeek-V4.1-Flash
How much GPU memory does DeepSeek-V4.1-Flash need?
About 1831.7 GB at 16-bit and 457.9 GB at 4-bit: the weights (763.2B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run DeepSeek-V4.1-Flash on?
At 16-bit, 8x MI325X from $16.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use DeepSeek-V4.1-Flash commercially?
Yes. DeepSeek-V4.1-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
What is DeepSeek-V4.1-Flash's context length?
1,048,576 tokens, from the maximum position embeddings in its published configuration.
Similar Models
We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…
We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…
deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction. The refusal direction was computed from 79 harmful vs. 79 benign instruction prompts (per-layer mean-difference of the collapsed residual stream, captured with the official reference implementation, tensor-parallel 4). Exactly 80 tensors were orthogonalized — for each of the 40 backbone layers: - layers.N.attn.wob.weight — attention output projection (writes into the residual stream) - layers.N.ffn.sharedexperts.w2.weight — shared-expert down projection Each weight W was edited as W ← W − r̂ (r̂ᵀ W) with r̂ the unit refusal direction of that layer, removing the model's ability to write the refusal…
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…
Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…