SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-V4-Flash-0731

by DeepSeek deepseek-ai/DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e.

Parameters304.2B
Context1,048,576
Weights166.9 GB
Licensemit
AccessOpen weights
Monthly Downloads4.3M

Runs On

What it takes to serve DeepSeek-V4-Flash-0731 (304.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 608.4 GB 730.0 GB 3x MI325X (256 GB)
Vultr
$6.00 4x MI300X $7.40 · 3x MI355X $7.77
8-bit 304.2 GB 365.0 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
4-bit 152.1 GB 182.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on DeepSeek-V4-Flash-0731

Only 6 of the 256 routed experts fire on any given token, but all 304.2B parameters have to be resident, and that sets the bill for DeepSeek-V4-Flash-0731. At 16-bit the working set is 730 GB: three MI325X cards with 256 GB each at $6.00 an hour. At 8-bit, 365 GB fits on two MI300X at $3.70; at 4-bit, 182.5 GB fits on one MI300X at $1.85. It generates text over a 1,048,576-token context.

SAVRN Index host prices run from $0.06 in and $0.18 out per million tokens at DeepInfra to $0.44 and $1.32 at Novita, more than a seven-fold spread, so price a run yourself. MIT keeps that simple: commercial use, modification and redistribution, provided the copyright and permission notice travels with the files. Decide how much of the million-token window you will run, since the memory figures are quoted on the weights, and read arXiv:2606.19348, the paper behind it.

Model Card

By DeepSeek, published under mit, revision 7872f01b1d1f.

Technical Report

Introduction

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.

Read the full model card (657 words)

Configuration

Architecture
DeepseekV4ForCausalLM
Context length (tokens)
1,048,576
Layers
43
Hidden size
4,096
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
256
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v4
Quantization
fp8

Identity and Version

Repository
deepseek-ai/DeepSeek-V4-Flash-0731
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
304.2B parameters
Languages
Not stated by the source
Revision
7872f01b1d1fe23eabc4c98b48bffcef5a386062
First published
2026-07-31
Last updated
2026-08-01

Files and Weights

74 files, 166.9 GB in total. The weights are 48 files totalling 166.9 GB in safetensors.

Weights48 files · 166.9 GB
Configuration14 files · 5.7 MB
Tokenizer2 files · 6.4 MB
Documentation4 files · 18.4 KB
Other5 files · 8.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights1.1 GB f3668ba4cccf
model-00002-of-00048.safetensorsWeights3.6 GB 77b26c939a0e
model-00003-of-00048.safetensorsWeights3.6 GB 412abf4c906f
model-00004-of-00048.safetensorsWeights3.6 GB 9610f56bc587
model-00005-of-00048.safetensorsWeights3.6 GB f87a5ac7b8be
model-00006-of-00048.safetensorsWeights3.6 GB 4a4f3764e3fc
model-00007-of-00048.safetensorsWeights3.6 GB df81bb80e27a
model-00008-of-00048.safetensorsWeights3.6 GB 224968d2b27f
model-00009-of-00048.safetensorsWeights3.6 GB 04d69ef1071f
model-00010-of-00048.safetensorsWeights3.6 GB 627145f4ebeb
model-00011-of-00048.safetensorsWeights3.6 GB e4b8e601dcbe
model-00012-of-00048.safetensorsWeights3.6 GB 64ed4e5f6126
model-00013-of-00048.safetensorsWeights3.6 GB 8dfe199d07c0
model-00014-of-00048.safetensorsWeights3.6 GB 45db2f540f82
model-00015-of-00048.safetensorsWeights3.6 GB 5810381a0f05
model-00016-of-00048.safetensorsWeights3.6 GB e0530b702477
model-00017-of-00048.safetensorsWeights3.6 GB ed1113024711
model-00018-of-00048.safetensorsWeights3.6 GB e393fea96da2
model-00019-of-00048.safetensorsWeights3.6 GB a74ca4d3e8e8
model-00020-of-00048.safetensorsWeights3.6 GB 9f556769926e
model-00021-of-00048.safetensorsWeights3.6 GB 1671cce7f90d
model-00022-of-00048.safetensorsWeights3.6 GB decd67a4bd97
model-00023-of-00048.safetensorsWeights3.6 GB c61a3e179cdb
model-00024-of-00048.safetensorsWeights3.6 GB fc27aeb42335
model-00025-of-00048.safetensorsWeights3.6 GB a66b6b8d5821
model-00026-of-00048.safetensorsWeights3.6 GB 657b89314fba
model-00027-of-00048.safetensorsWeights3.6 GB fb01f21a0da0
model-00028-of-00048.safetensorsWeights3.6 GB b2fd5cbbb639
model-00029-of-00048.safetensorsWeights3.6 GB 9ec2fdf90027
model-00030-of-00048.safetensorsWeights3.6 GB 9ed3c317bf96
model-00031-of-00048.safetensorsWeights3.6 GB d5078c3fca3e
model-00032-of-00048.safetensorsWeights3.6 GB 163653848f00
model-00033-of-00048.safetensorsWeights3.6 GB f2cffd43f2a5
model-00034-of-00048.safetensorsWeights3.6 GB 0f9494512147
model-00035-of-00048.safetensorsWeights3.6 GB 9cb6a316989f
model-00036-of-00048.safetensorsWeights3.6 GB 7e6761421fe9
model-00037-of-00048.safetensorsWeights3.6 GB a59d662f1143
model-00038-of-00048.safetensorsWeights3.6 GB 137fa617a74b
model-00039-of-00048.safetensorsWeights3.6 GB a29af1aa519d
model-00040-of-00048.safetensorsWeights3.6 GB 8bc93d8a7d19
model-00041-of-00048.safetensorsWeights3.6 GB fd312e7fdd6c
model-00042-of-00048.safetensorsWeights3.6 GB 4d19bf368083
model-00043-of-00048.safetensorsWeights3.6 GB b7103842ceb7
model-00044-of-00048.safetensorsWeights3.6 GB 422d3889fa20
model-00045-of-00048.safetensorsWeights1.1 GB a5be6aed7b84
model-00046-of-00048.safetensorsWeights3.6 GB 5db924ca907e
model-00047-of-00048.safetensorsWeights3.6 GB 62816173f9f6
model-00048-of-00048.safetensorsWeights3.7 GB cc43742bd24a
config.jsonConfiguration1.9 KB
encoding/encoding_dsv4.pyConfiguration29.0 KB
encoding/test_encoding_dsv4.pyConfiguration3.7 KB
encoding/tests/test_input_1.jsonConfiguration2.7 KB
encoding/tests/test_input_2.jsonConfiguration526 B
encoding/tests/test_input_3.jsonConfiguration4.5 KB
encoding/tests/test_input_4.jsonConfiguration2.7 KB
generation_config.jsonConfiguration170 B
inference/config.jsonConfiguration1.2 KB
inference/convert.pyConfiguration6.6 KB
inference/generate.pyConfiguration5.8 KB
inference/kernel.pyConfiguration22.2 KB
inference/model.pyConfiguration45.1 KB
model.safetensors.index.jsonConfiguration5.6 MB
LICENSEDocumentation1.1 KB
README.mdDocumentation7.2 KB
encoding/README.mdDocumentation9.2 KB
inference/README.mdDocumentation951 B
encoding/tests/test_output_1.txtOther2.4 KB
encoding/tests/test_output_2.txtOther342 B
encoding/tests/test_output_3.txtOther3.3 KB
encoding/tests/test_output_4.txtOther2.6 KB
inference/requirements.txtOther92 B
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer6.4 MB
tokenizer_config.jsonTokenizer801 B

License and Download

License
mit
Access
Open weights, no gate
Download size
166.9 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
datacurve/deep-swe Task deep_sweMetric deep_sweComparison conditions not established 54.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-03
harborframework/terminal-bench-2.1 Task terminalbench_2_1Metric terminalbench_2_1Setup DeepSeek Harness (minimal mode), max reasoning effort, temperature=1.0, top_p=0.95.Comparison conditions not established 82.7 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-01
hkust-nlp/Toolathlon Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established 70.3 deepseek-ai/DeepSeek-V4-Flash-0731 model card
Reported by a third party
Evaluated revision not stated 2026-08-01

Memory Requirements

PrecisionWeights in memory
As published166.9 GB
16-bit608.4 GB
8-bit304.2 GB
4-bit152.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
Baseten$0.13 / $0.26input / output, per million tokensSep 18, 2026
DeepInfra$0.06 / $0.18input / output, per million tokensSep 18, 2026
Fireworks$0.22 / $0.66input / output, per million tokensSep 18, 2026
Novita$0.44 / $1.32input / output, per million tokensSep 18, 2026
Together AI$0.14 / $0.28input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Compare DeepSeek-V4-Flash-0731

Questions About DeepSeek-V4-Flash-0731

How much GPU memory does DeepSeek-V4-Flash-0731 need?

About 730 GB at 16-bit and 182.5 GB at 4-bit: the weights (304.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V4-Flash-0731 on?

At 16-bit, 3x MI325X from $6.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V4-Flash-0731 commercially?

Yes. DeepSeek-V4-Flash-0731 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V4-Flash-0731's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

DeepSeek-V4-Flash

DeepSeek

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache…

Open weights mit 290.9B parameters 1,048,576 tokens transformers

Model · Text generation

MiniMax-M2.7

MiniMax

Join Our WeChat Discord community. MiniMax Agent API CLI MiniMax Website Hugging Face GitHub ModelScope LICENSE MiniMax-M2.7 is our first model deeply participating in its own evolution. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search. For more details, see our blog post. M2.7 initiates a cycle of model self-evolution: during development, we let the model update its own memory, build dozens of complex skills for RL experiments, and improve its own learning process based on experiment results. An internal version of M2.7 autonomously optimized a programming scaffold over…

Open weights other 228.7B parameters 204,800 tokens transformers

Model · Text generation

GLM-5.2-NVFP4

NVIDIA

The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM-5.2 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial or non-commercial use. GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG…

Open weights mit 381B parameters 1,048,576 tokens Model Optimizer

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

Model · Text generation

Qwen3-Coder-480B-A35B-Instruct-FP8

Qwen

Today, we're announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we're excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct. featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks, achieving results comparable to Claude Sonnet. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.…

Open weights apache-2.0 480.2B parameters 262,144 tokens transformers

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers