SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-V4-Flash

by DeepSeek deepseek-ai/DeepSeek-V4-Flash

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context…

Parameters290.9B
Context1,048,576
Weights159.6 GB
Licensemit
AccessOpen weights
Monthly Downloads1.6M

Runs On

What it takes to serve DeepSeek-V4-Flash (290.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 581.9 GB 698.3 GB 3x MI325X (256 GB)
Vultr
$6.00 4x MI300X $7.40 · 3x MI355X $7.77
8-bit 290.9 GB 349.1 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
4-bit 145.5 GB 174.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By DeepSeek, published under mit, revision 60d8d70770c6.

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Technical Report

Introduction

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:

  1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2.
  2. Manifold-Constrained Hyper-Connections (mHC): We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity.
  3. Muon Optimizer: We employ the Muon optimizer for faster convergence and greater training stability.

Read the full model card (1,925 words)

Configuration

Architecture
DeepseekV4ForCausalLM
Context length (tokens)
1,048,576
Layers
43
Hidden size
4,096
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
256
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v4
Quantization
fp8

Identity and Version

Repository
deepseek-ai/DeepSeek-V4-Flash
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
290.9B parameters
Languages
Not stated by the source
Revision
60d8d70770c6776ff598c94bb586a859a38244f1
First published
2026-04-22
Last updated
2026-06-22

Files and Weights

73 files, 159.6 GB in total. The weights are 46 files totalling 159.6 GB in safetensors.

Weights46 files · 159.6 GB
Configuration14 files · 5.5 MB
Tokenizer2 files · 6.4 MB
Documentation4 files · 23.3 KB
Other6 files · 1.0 MB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00046.safetensorsWeights1.1 GB 517658661390
model-00002-of-00046.safetensorsWeights3.6 GB f04048189d3b
model-00003-of-00046.safetensorsWeights3.6 GB df5f80b9b4ca
model-00004-of-00046.safetensorsWeights3.6 GB 948250b46a6f
model-00005-of-00046.safetensorsWeights3.6 GB 9fda158bc636
model-00006-of-00046.safetensorsWeights3.6 GB 51a65e6d9d0c
model-00007-of-00046.safetensorsWeights3.6 GB 2d782c46d6d2
model-00008-of-00046.safetensorsWeights3.6 GB b7d9d8d8932e
model-00009-of-00046.safetensorsWeights3.6 GB 3197a42d282a
model-00010-of-00046.safetensorsWeights3.6 GB c9cef4200444
model-00011-of-00046.safetensorsWeights3.6 GB 916b0b34b713
model-00012-of-00046.safetensorsWeights3.6 GB 05739c7d91a3
model-00013-of-00046.safetensorsWeights3.6 GB 47c5e416b60b
model-00014-of-00046.safetensorsWeights3.6 GB c881e2671ab4
model-00015-of-00046.safetensorsWeights3.6 GB cb3daae8e465
model-00016-of-00046.safetensorsWeights3.6 GB 9eba661fba31
model-00017-of-00046.safetensorsWeights3.6 GB 0c36cbc026c5
model-00018-of-00046.safetensorsWeights3.6 GB 1298a07452a4
model-00019-of-00046.safetensorsWeights3.6 GB d3687748aff7
model-00020-of-00046.safetensorsWeights3.6 GB 906c652f3c36
model-00021-of-00046.safetensorsWeights3.6 GB f270bf4d0f00
model-00022-of-00046.safetensorsWeights3.6 GB d02261b8f1c8
model-00023-of-00046.safetensorsWeights3.6 GB 69fab8bfa1cd
model-00024-of-00046.safetensorsWeights3.6 GB baba23c06a7b
model-00025-of-00046.safetensorsWeights3.6 GB 085b7736ebe3
model-00026-of-00046.safetensorsWeights3.6 GB fdde6791ab71
model-00027-of-00046.safetensorsWeights3.6 GB 2f207b9aef9c
model-00028-of-00046.safetensorsWeights3.6 GB 2cc519b5a03a
model-00029-of-00046.safetensorsWeights3.6 GB d10bf34c789f
model-00030-of-00046.safetensorsWeights3.6 GB 0f1c471fa9d9
model-00031-of-00046.safetensorsWeights3.6 GB beafa59d64fa
model-00032-of-00046.safetensorsWeights3.6 GB 5c6b2934d87a
model-00033-of-00046.safetensorsWeights3.6 GB c05f917e873d
model-00034-of-00046.safetensorsWeights3.6 GB 666b77201ec6
model-00035-of-00046.safetensorsWeights3.6 GB 5e6b9a54fd14
model-00036-of-00046.safetensorsWeights3.6 GB ab72ad9d171f
model-00037-of-00046.safetensorsWeights3.6 GB 93d68bcfc36f
model-00038-of-00046.safetensorsWeights3.6 GB 809fb799edcf
model-00039-of-00046.safetensorsWeights3.6 GB 49dba2489174
model-00040-of-00046.safetensorsWeights3.6 GB 09a7b8b6957f
model-00041-of-00046.safetensorsWeights3.6 GB a564ac6c6cc7
model-00042-of-00046.safetensorsWeights3.6 GB bd3f5b898b04
model-00043-of-00046.safetensorsWeights3.6 GB 85a414c7991c
model-00044-of-00046.safetensorsWeights3.6 GB 438b052b8a2d
model-00045-of-00046.safetensorsWeights1.1 GB 9a0fd242134e
model-00046-of-00046.safetensorsWeights3.6 GB f58f722893a6
config.jsonConfiguration1.7 KB
encoding/encoding_dsv4.pyConfiguration27.9 KB
encoding/test_encoding_dsv4.pyConfiguration3.7 KB
encoding/tests/test_input_1.jsonConfiguration2.7 KB
encoding/tests/test_input_2.jsonConfiguration526 B
encoding/tests/test_input_3.jsonConfiguration4.5 KB
encoding/tests/test_input_4.jsonConfiguration2.7 KB
generation_config.jsonConfiguration170 B
inference/config.jsonConfiguration991 B
inference/convert.pyConfiguration7.1 KB
inference/generate.pyConfiguration6.3 KB
inference/kernel.pyConfiguration22.2 KB
inference/model.pyConfiguration38.6 KB
model.safetensors.index.jsonConfiguration5.4 MB
LICENSEDocumentation1.1 KB
README.mdDocumentation13.1 KB
encoding/README.mdDocumentation8.1 KB
inference/README.mdDocumentation951 B
assets/dsv4_performance.pngOther1.0 MB 8fd472981a4c
encoding/tests/test_output_1.txtOther2.4 KB
encoding/tests/test_output_2.txtOther342 B
encoding/tests/test_output_3.txtOther3.3 KB
encoding/tests/test_output_4.txtOther2.6 KB
inference/requirements.txtOther92 B
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer6.4 MB
tokenizer_config.jsonTokenizer801 B

License and Download

License
mit
Access
Open weights, no gate
Download size
159.6 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 88.1 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-24
IntelligenceLab/Long-Horizon-Terminal-Bench Task lhtb_solvedMetric lhtb_solvedSetup 2/46 tasks solved at reward >= 0.95, estimated at the leaderboard's 90-minute budget from interim verifier checkpoints of a 3-hour terminus-2 run; NOT a real 90-minute run. Range 2-3/46 (riscv-core-debug unresolved). The same run scores 5/46 at its full 3-hour budget. Official LHTB Harbor harness.Comparison conditions not established 2 LHTB run artifacts
Reported by a third party
Evaluated revision not stated 2026-08-08
SWE-bench/SWE-bench_Multilingual Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established 73.3 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-10
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 79 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-24
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 86.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-24
benchflow/skillsbench Task skillsbench_v1_1Metric skillsbench_v1_1Setup with-skills; BenchFlow harness; OpenHands agent; 87 tasks x 3 trials; full 261/261 coverageComparison conditions not established 44.7 SkillsBench v1.1 official leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-11
cais/hle Task hleMetric hleComparison conditions not established 34.8 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-24
claw-eval/Claw-Eval Task generalMetric generalSetup Pass³% | N=3 | 161 tasksComparison conditions not established 57.8 Claw-Eval Leaderboard
Reported by a third party
Evaluated revision not stated 2026-04-23
claw-eval/Claw-Eval Task multi_turnMetric multi_turnSetup Pass³% | N=3 | 38 tasksComparison conditions not established 57.9 Claw-Eval Leaderboard
Reported by a third party
Evaluated revision not stated 2026-04-23
harborframework/terminal-bench-2.0 Task terminalbench_2Metric terminalbench_2Comparison conditions not established 56.9 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-24
joelniklaus/LEXam-hard Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established 38.19 SwissLegalEvals per-sample details (lighteval)
Reported by a third party
Evaluated revision not stated 2026-06-12

Memory Requirements

PrecisionWeights in memory
As published159.6 GB
16-bit581.9 GB
8-bit290.9 GB
4-bit145.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.09 / $0.18input / output, per million tokensSep 18, 2026
Novita$0.14 / $0.28input / output, per million tokensSep 18, 2026
Vultr$0.30 / $1.00input / output, per million tokensSep 10, 2026

From the SAVRN Index.

Built on This Model

Questions About DeepSeek-V4-Flash

How much GPU memory does DeepSeek-V4-Flash need?

About 698.3 GB at 16-bit and 174.6 GB at 4-bit: the weights (290.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V4-Flash on?

At 16-bit, 3x MI325X from $6.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V4-Flash commercially?

Yes. DeepSeek-V4-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V4-Flash's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

DeepSeek-V4-Flash-0731

DeepSeek

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. 1. For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort…

Open weights mit 304.2B parameters 1,048,576 tokens transformers

Model · Text generation

MiniMax-M2.7

MiniMax

Join Our WeChat Discord community. MiniMax Agent API CLI MiniMax Website Hugging Face GitHub ModelScope LICENSE MiniMax-M2.7 is our first model deeply participating in its own evolution. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search. For more details, see our blog post. M2.7 initiates a cycle of model self-evolution: during development, we let the model update its own memory, build dozens of complex skills for RL experiments, and improve its own learning process based on experiment results. An internal version of M2.7 autonomously optimized a programming scaffold over…

Open weights other 228.7B parameters 204,800 tokens transformers

Model · Text generation

GLM-5.2-NVFP4

NVIDIA

The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM-5.2 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial or non-commercial use. GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG…

Open weights mit 381B parameters 1,048,576 tokens Model Optimizer

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

Model · Text generation

gpt-oss-120b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 116.8B parameters 131,072 tokens transformers