SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-V4-Flash-DSpark

by DeepSeek deepseek-ai/DeepSeek-V4-Flash-DSpark

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder.

Parameters165.3B
Context1,048,576
Weights166.9 GB
Licensemit
AccessOpen weights
Monthly Downloads1M

Runs On

What it takes to serve DeepSeek-V4-Flash-DSpark (165.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 330.5 GB 396.6 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55
8-bit 165.3 GB 198.3 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70
4-bit 82.6 GB 99.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on DeepSeek-V4-Flash-DSpark

The publisher says it outright: DSpark is not a new model. It is the DeepSeek-V4-Flash checkpoint with a speculative decoding module attached, so you are evaluating a serving path, not fresh weights. Those weights run 165.3 billion parameters across 256 routed experts with 6 active per token. Precision picks the hardware: 16-bit needs 396.6 GB and two MI325X cards at $4.00 per hour, 8-bit needs 198.3 GB on one MI325X at $2.00, and 4-bit needs 99.2 GB on one MI300X at $1.85.

MIT covers it: commercial use, modification and redistribution with the notices kept. Two checks before buying cards. The memory figures are the floor; a request that uses the 1,048,576-token context adds cache on top, sized from the single key/value head at 512 dimensions. And the speculative decoding path is the point of this release, so confirm your serving stack runs the publisher's inference example.

Model Card

By DeepSeek, published under mit, revision 62af8fffb2f7.

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Technical Report

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec

Introduction

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:

Read the full model card (2,036 words)

Configuration

Architecture
DeepseekV4ForCausalLM
Context length (tokens)
1,048,576
Layers
43
Hidden size
4,096
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
256
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v4
Quantization
fp8

Identity and Version

Repository
deepseek-ai/DeepSeek-V4-Flash-DSpark
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
165.3B parameters
Languages
Not stated by the source
Revision
62af8fffb2f7030cac4de2f0169f5b8d1101b646
First published
2026-06-27
Last updated
2026-07-04

Files and Weights

74 files, 166.9 GB in total. The weights are 48 files totalling 166.9 GB in safetensors.

Weights48 files · 166.9 GB
Configuration14 files · 5.7 MB
Tokenizer2 files · 6.4 MB
Documentation4 files · 24.5 KB
Other5 files · 8.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights1.1 GB 517658661390
model-00002-of-00048.safetensorsWeights3.6 GB 0821952bc12c
model-00003-of-00048.safetensorsWeights3.6 GB 2bd2b020f81f
model-00004-of-00048.safetensorsWeights3.6 GB e6286ec182a2
model-00005-of-00048.safetensorsWeights3.6 GB 021dd40791d0
model-00006-of-00048.safetensorsWeights3.6 GB 3d96bfc9686d
model-00007-of-00048.safetensorsWeights3.6 GB 4481b0c86e9f
model-00008-of-00048.safetensorsWeights3.6 GB e19559e75d35
model-00009-of-00048.safetensorsWeights3.6 GB ab44718d102e
model-00010-of-00048.safetensorsWeights3.6 GB 743c16a35f1c
model-00011-of-00048.safetensorsWeights3.6 GB 4fc2c700363b
model-00012-of-00048.safetensorsWeights3.6 GB 5214de3c6ca8
model-00013-of-00048.safetensorsWeights3.6 GB d953f118488b
model-00014-of-00048.safetensorsWeights3.6 GB 40c97b77151b
model-00015-of-00048.safetensorsWeights3.6 GB 136cfaa14f36
model-00016-of-00048.safetensorsWeights3.6 GB 6881f71ac5da
model-00017-of-00048.safetensorsWeights3.6 GB a1949b35de69
model-00018-of-00048.safetensorsWeights3.6 GB 8288d4462015
model-00019-of-00048.safetensorsWeights3.6 GB 25a4d2a2a4a0
model-00020-of-00048.safetensorsWeights3.6 GB 0016a46d9f85
model-00021-of-00048.safetensorsWeights3.6 GB 3f342f4120b9
model-00022-of-00048.safetensorsWeights3.6 GB 0bf9870e1d44
model-00023-of-00048.safetensorsWeights3.6 GB 215f51b475b8
model-00024-of-00048.safetensorsWeights3.6 GB a11f085685e7
model-00025-of-00048.safetensorsWeights3.6 GB 1d2a9ec82956
model-00026-of-00048.safetensorsWeights3.6 GB aaec7eb32b16
model-00027-of-00048.safetensorsWeights3.6 GB 63069bb430af
model-00028-of-00048.safetensorsWeights3.6 GB 2a804253c04f
model-00029-of-00048.safetensorsWeights3.6 GB 68dfa669a272
model-00030-of-00048.safetensorsWeights3.6 GB dd8433e27dff
model-00031-of-00048.safetensorsWeights3.6 GB 6e7796ad3b05
model-00032-of-00048.safetensorsWeights3.6 GB 86d3f021bec5
model-00033-of-00048.safetensorsWeights3.6 GB 35a92fa5a3cc
model-00034-of-00048.safetensorsWeights3.6 GB 368896471e8e
model-00035-of-00048.safetensorsWeights3.6 GB 85c9a9af1ec5
model-00036-of-00048.safetensorsWeights3.6 GB 0e22e7f41331
model-00037-of-00048.safetensorsWeights3.6 GB 47f41e12ce98
model-00038-of-00048.safetensorsWeights3.6 GB cdbb8c36da34
model-00039-of-00048.safetensorsWeights3.6 GB 03f3cc8ab1a7
model-00040-of-00048.safetensorsWeights3.6 GB 3124105d3834
model-00041-of-00048.safetensorsWeights3.6 GB 4eda737587c4
model-00042-of-00048.safetensorsWeights3.6 GB 9ac23e2e32c4
model-00043-of-00048.safetensorsWeights3.6 GB 8c629a864a9b
model-00044-of-00048.safetensorsWeights3.6 GB 745476d81cb1
model-00045-of-00048.safetensorsWeights1.1 GB 9a0fd242134e
model-00046-of-00048.safetensorsWeights3.6 GB 14810f274692
model-00047-of-00048.safetensorsWeights3.6 GB 7a44164698d9
model-00048-of-00048.safetensorsWeights3.7 GB a0bbb24f36d2
config.jsonConfiguration1.9 KB
encoding/encoding_dsv4.pyConfiguration27.9 KB
encoding/test_encoding_dsv4.pyConfiguration3.7 KB
encoding/tests/test_input_1.jsonConfiguration2.7 KB
encoding/tests/test_input_2.jsonConfiguration526 B
encoding/tests/test_input_3.jsonConfiguration4.5 KB
encoding/tests/test_input_4.jsonConfiguration2.7 KB
generation_config.jsonConfiguration170 B
inference/config.jsonConfiguration1.2 KB
inference/convert.pyConfiguration6.6 KB
inference/generate.pyConfiguration5.8 KB
inference/kernel.pyConfiguration22.2 KB
inference/model.pyConfiguration45.1 KB
model.safetensors.index.jsonConfiguration5.6 MB
LICENSEDocumentation1.1 KB
README.mdDocumentation14.4 KB
encoding/README.mdDocumentation8.1 KB
inference/README.mdDocumentation951 B
encoding/tests/test_output_1.txtOther2.4 KB
encoding/tests/test_output_2.txtOther342 B
encoding/tests/test_output_3.txtOther3.3 KB
encoding/tests/test_output_4.txtOther2.6 KB
inference/requirements.txtOther92 B
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer6.4 MB
tokenizer_config.jsonTokenizer801 B

License and Download

License
mit
Access
Open weights, no gate
Download size
166.9 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published166.9 GB
16-bit330.5 GB
8-bit165.3 GB
4-bit82.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare DeepSeek-V4-Flash-DSpark

Questions About DeepSeek-V4-Flash-DSpark

How much GPU memory does DeepSeek-V4-Flash-DSpark need?

About 396.6 GB at 16-bit and 99.2 GB at 4-bit: the weights (165.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V4-Flash-DSpark on?

At 16-bit, 2x MI325X from $4.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V4-Flash-DSpark commercially?

Yes. DeepSeek-V4-Flash-DSpark is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V4-Flash-DSpark's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

Model · Text generation

gpt-oss-120b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 116.8B parameters 131,072 tokens transformers

Model · Text generation

MiniMax-M2.7

MiniMax

Join Our WeChat Discord community. MiniMax Agent API CLI MiniMax Website Hugging Face GitHub ModelScope LICENSE MiniMax-M2.7 is our first model deeply participating in its own evolution. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search. For more details, see our blog post. M2.7 initiates a cycle of model self-evolution: during development, we let the model update its own memory, build dozens of complex skills for RL experiments, and improve its own learning process based on experiment results. An internal version of M2.7 autonomously optimized a programming scaffold over…

Open weights other 228.7B parameters 204,800 tokens transformers

Model · Text generation

Qwen3-Coder-Next-FP8

Qwen

Today, we're announcing Qwen3-Coder-Next-FP8, an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements: Qwen3-Coder-Next-FP8 has the following features: NOTE: This model supports only non-thinking mode and does not generate blocks in its output. Meanwhile, specifying enablethinking=False is no longer required. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. We advise you to use the latest version of transformers. The following contains a code snippet illustrating how to use the model generate content based on…

Open weights apache-2.0 79.7B parameters 262,144 tokens transformers

Model · Text generation

Qwen-72B

Qwen

通义千问-72B(Qwen-72B)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B的仓库。 通义千问-72B(Qwen-72B)主要有以下特点: 1. 大规模高质量训练语料:使用超过3万亿tokens的数据进行预训练,包含高质量中、英、多语言、代码、数学等数据,涵盖通用及专业领域的训练语料。通过大量对比实验对预训练语料分布进行了优化。 2. 强大的性能:Qwen-72B在多个中英文下游评测任务上(涵盖常识推理、代码、数学、翻译等),效果显著超越现有的开源模型。具体评测结果请详见下文。 3. 覆盖更全面的词表:相比目前以中英词表为主的开源模型,Qwen-72B使用了约15万大小的词表。该词表对多语言更加友好,方便用户在不扩展词表的情况下对部分语种进行能力增强和扩展。 4. 较长的上下文支持:Qwen-72B支持32k的上下文长度。 Qwen-72B is the 72B-parameter version of the large language model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-72B is a Transformer-based large…

Open weights other 72.3B parameters 32,768 tokens transformers

Model · Text generation

Llama-3.3-70B-Instruct

Meta Llama

The Meta Llama 3.3 multilingual large language model (LLM) is an instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.3 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and…

Access requested at publisher llama3.3 70.6B parameters transformers