SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-V4-Pro-0813

by SAIFI INDUSTRIES SAIFIINDUSTRIES/DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813 is an open-weight model for text generation from SAIFI INDUSTRIES, released under MIT License. It has 1.7T parameters and a 1,048,576-token context. Its published files total 892.8 GB.

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments.

Parameters1.7T
Context1,048,576
Weights892.7 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve DeepSeek-V4-Pro-0813 (1.7T parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3301.0 GB 3961.2 GB More than one server of any accelerator the SAVRN Index prices.
8-bit 1650.5 GB 1980.6 GB 8x MI325X (256 GB)
Vultr
$16.00 7x MI355X $18.13 · 8x B300 $52.80
4-bit 825.2 GB 990.3 GB 4x MI325X (256 GB)
Vultr
$8.00 4x MI355X $10.36 · 6x MI300X $11.10

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

DeepSeek-V4-Pro-0813 on every accelerator the SAVRN Index prices, at every precision

Model Card

By SAIFI INDUSTRIES, published under mit, revision 74ba95efd0d5.

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached. DeepSeek-V4-Pro-0813 outperforms DeepSeek-V4-Pro (Preview) on the benchmarks listed below, and is broadly competitive with the strongest proprietary models available. 1. For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning…

Read SAIFI INDUSTRIES's full model card

Technical Report

Introduction

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.

DeepSeek-V4-Pro-0813 outperforms DeepSeek-V4-Pro (Preview) on the benchmarks listed below, and is broadly competitive with the strongest proprietary models available.

| Benchmark | DeepSeek-V4-Pro-0813 | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Pro (Preview) | DeepSeek-V4-Flash (Preview) | GLM-5.2 | Kimi K3 | Opus-4.8 | Fable-5 (w/ fallback) | | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | HLE (wo / w tools) | 42.7 / 60.0 | 37.8 / 51.5 | 37.7 / 48.2 | 34.8 / 45.1 | 40.5 / 54.7 | 43.5 / 56.0 | 49.8 / 57.9 | 53.3 / 63.0 | | Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 | | NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - | | Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 | | DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 | | Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 | | Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - | | AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 | | DSBench-FullStack † | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 | | DSBench-Hard † | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 |

Notes:

  1. For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95.
  2. † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.

Chat Template

This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.

The reasoning_effort parameter now supports three levels — low, high, and max — which control how much deliberation the model spends before answering.

A brief example:

from encoding_dsv4 import encode_messages, parse_message_from_completion_text

messages = [
    {"role": "user", "content": "hello"},
    {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
    {"role": "user", "content": "1+1=?"}
]

# messages -> string
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")

# string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro-0813")
tokens = tokenizer.encode(prompt)

How to Run with vLLM

DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command:

--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'

For example, the command below serves the model with vLLM on a single 4×GB300 node. See the vLLM recipe for detailed instructions and other hardware configurations.

vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 \
  --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
  --data-parallel-size 4 --enable-expert-parallel \
  --moe-backend deep_gemm_mega_moe \
  --attention-config '{"use_fp4_indexer_cache": true}' \
  --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'

How to Run with SGLang

Enable DSpark with --speculative-algorithm DSPARK and do not set a separate --speculative-draft-model-path as the target and draft weights therefore come from the same checkpoint. See the SGLang cookbook for detailed instructions, benchmarks and other hardwares configurations.

sglang serve \
  --trust-remote-code \
  --model-path deepseek-ai/DeepSeek-V4-Pro-0813 \
  --tp 4 \
  --moe-runner-backend flashinfer_mxfp4 \
  --speculative-algorithm DSPARK \
  --mem-fraction-static 0.90 \
  --chunked-prefill-size 4096 \
  --swa-full-tokens-ratio 0.1 \

How to Run Locally

Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.

For local deployment, we recommend setting the sampling parameters to temperature = 1.0, with top_p = 0.95 for agentic scenarios and top_p = 1.0 otherwise. For the high and max reasoning effort levels, we recommend a maximum output length of 384K tokens.

License

This repository and the model weights are licensed under the MIT License.

Citation

@misc{deepseekai2026deepseekv4,
      title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
      author={DeepSeek-AI},
      year={2026},
}

Contact

If you have any questions, please raise an issue or contact us at [email protected].

Configuration

Architecture
DeepseekV4ForCausalLM
Context length (tokens)
1,048,576
Layers
61
Hidden size
7,168
Attention heads
128
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
384
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v4
Quantization
fp8

Identity and Version

Repository
SAIFIINDUSTRIES/DeepSeek-V4-Pro-0813
Publisher
SAIFI INDUSTRIES
Task
Text generation
Modality
Text
Library
transformers
Parameters
1.7T parameters
Languages
Not stated by the source
Revision
74ba95efd0d516f8c94b7455848b1f0881cb99d1
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

92 files, 892.8 GB in total. The weights are 66 files totalling 892.7 GB in safetensors.

Weights66 files · 892.7 GB
Configuration14 files · 11.8 MB
Tokenizer2 files · 6.4 MB
Documentation4 files · 18.7 KB
Other5 files · 8.7 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00066.safetensorsWeights1.9 GB 74ff60e30429
model-00002-of-00066.safetensorsWeights13.9 GB d7be184a5582
model-00003-of-00066.safetensorsWeights13.9 GB fc9291c02445
model-00004-of-00066.safetensorsWeights13.9 GB 6f2c6511e5cc
model-00005-of-00066.safetensorsWeights13.9 GB 453aeaf23b4c
model-00006-of-00066.safetensorsWeights13.9 GB 9a1ea3ba51ac
model-00007-of-00066.safetensorsWeights13.9 GB 1ac789c43cbf
model-00008-of-00066.safetensorsWeights13.9 GB b2e31dda0b24
model-00009-of-00066.safetensorsWeights13.9 GB b2efb301a7e2
model-00010-of-00066.safetensorsWeights13.9 GB 9f0d7228d9f9
model-00011-of-00066.safetensorsWeights13.9 GB 9ff42cf9b641
model-00012-of-00066.safetensorsWeights13.9 GB c50346bb61a3
model-00013-of-00066.safetensorsWeights13.9 GB e4fd1c5c65e0
model-00014-of-00066.safetensorsWeights13.9 GB a72fd0426270
model-00015-of-00066.safetensorsWeights13.9 GB ca601d33d249
model-00016-of-00066.safetensorsWeights13.9 GB f3ce6de8fbb5
model-00017-of-00066.safetensorsWeights13.9 GB 4af598f3672a
model-00018-of-00066.safetensorsWeights13.9 GB 0313cffc575c
model-00019-of-00066.safetensorsWeights13.9 GB 5ea603dbdab6
model-00020-of-00066.safetensorsWeights13.9 GB c3bb54f871d8
model-00021-of-00066.safetensorsWeights13.9 GB 516167a84b4b
model-00022-of-00066.safetensorsWeights13.9 GB df3975926bf3
model-00023-of-00066.safetensorsWeights13.9 GB e444e6d51548
model-00024-of-00066.safetensorsWeights13.9 GB 8690577cde63
model-00025-of-00066.safetensorsWeights13.9 GB 170edd2db317
model-00026-of-00066.safetensorsWeights13.9 GB b3dcaa3372a4
model-00027-of-00066.safetensorsWeights13.9 GB 50d385e1e247
model-00028-of-00066.safetensorsWeights13.9 GB 0449120425ee
model-00029-of-00066.safetensorsWeights13.9 GB 37e8b8bcb439
model-00030-of-00066.safetensorsWeights13.9 GB f9bc11eef0b4
model-00031-of-00066.safetensorsWeights13.9 GB f115720429d7
model-00032-of-00066.safetensorsWeights13.9 GB 7ac9402f531e
model-00033-of-00066.safetensorsWeights13.9 GB dcb6e9370ec5
model-00034-of-00066.safetensorsWeights13.9 GB 607f907a970a
model-00035-of-00066.safetensorsWeights13.9 GB ca4454ca96c5
model-00036-of-00066.safetensorsWeights13.9 GB 4ac44a65a92d
model-00037-of-00066.safetensorsWeights13.9 GB 6aad5ecd60cd
model-00038-of-00066.safetensorsWeights13.9 GB b8e9376eeba0
model-00039-of-00066.safetensorsWeights13.9 GB 5182b6d508d4
model-00040-of-00066.safetensorsWeights13.9 GB dc91a11234ee
model-00041-of-00066.safetensorsWeights13.9 GB 1b446abc23a0
model-00042-of-00066.safetensorsWeights13.9 GB 5ec38fcfe390
model-00043-of-00066.safetensorsWeights13.9 GB a9279341f553
model-00044-of-00066.safetensorsWeights13.9 GB ef67de8c818b
model-00045-of-00066.safetensorsWeights13.9 GB 5830f22254b5
model-00046-of-00066.safetensorsWeights13.9 GB 22ae4fcc55b9
model-00047-of-00066.safetensorsWeights13.9 GB 30890437fc59
model-00048-of-00066.safetensorsWeights13.9 GB 877e6060d9c1
model-00049-of-00066.safetensorsWeights13.9 GB 0527cff3c52c
model-00050-of-00066.safetensorsWeights13.9 GB 52a94fd51d93
model-00051-of-00066.safetensorsWeights13.9 GB 50044649fa7d
model-00052-of-00066.safetensorsWeights13.9 GB 086f263f3f49
model-00053-of-00066.safetensorsWeights13.9 GB 5e0264fd0c71
model-00054-of-00066.safetensorsWeights13.9 GB 3f3adcbb18ea
model-00055-of-00066.safetensorsWeights13.9 GB 5d4ab723dfab
model-00056-of-00066.safetensorsWeights13.9 GB 409be8499d63
model-00057-of-00066.safetensorsWeights13.9 GB 75146b06c0eb
model-00058-of-00066.safetensorsWeights13.9 GB c3bbbde96d32
model-00059-of-00066.safetensorsWeights13.9 GB 2f5b62cd3111
model-00060-of-00066.safetensorsWeights13.9 GB 0e4ac232d193
model-00061-of-00066.safetensorsWeights13.9 GB dc3ef841f627
model-00062-of-00066.safetensorsWeights13.9 GB c8d8ecc7bc03
model-00063-of-00066.safetensorsWeights1.9 GB 56193a52a6fc
model-00064-of-00066.safetensorsWeights14.0 GB 3869820fbd71
model-00065-of-00066.safetensorsWeights13.9 GB 09268dadf4bb
model-00066-of-00066.safetensorsWeights14.1 GB 40d4bf1620bd
config.jsonConfiguration2.0 KB —
encoding/encoding_dsv4.pyConfiguration29.0 KB —
encoding/test_encoding_dsv4.pyConfiguration3.7 KB —
encoding/tests/test_input_1.jsonConfiguration2.7 KB —
encoding/tests/test_input_2.jsonConfiguration526 B —
encoding/tests/test_input_3.jsonConfiguration4.5 KB —
encoding/tests/test_input_4.jsonConfiguration2.7 KB —
generation_config.jsonConfiguration170 B —
inference/config.jsonConfiguration1.2 KB —
inference/convert.pyConfiguration6.6 KB —
inference/generate.pyConfiguration5.8 KB —
inference/kernel.pyConfiguration22.2 KB —
inference/model.pyConfiguration45.1 KB —
model.safetensors.index.jsonConfiguration11.7 MB 2de2ac1e4313
LICENSEDocumentation1.1 KB —
README.mdDocumentation7.5 KB —
encoding/README.mdDocumentation9.2 KB —
inference/README.mdDocumentation951 B —
encoding/tests/test_output_1.txtOther2.4 KB —
encoding/tests/test_output_2.txtOther342 B —
encoding/tests/test_output_3.txtOther3.3 KB —
encoding/tests/test_output_4.txtOther2.6 KB —
inference/requirements.txtOther92 B —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer6.4 MB —
tokenizer_config.jsonTokenizer801 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
892.7 GB
Download from SAIFI INDUSTRIES

Released by SAIFI INDUSTRIES through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published892.7 GB
16-bit3301.0 GB
8-bit1650.5 GB
4-bit825.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DeepSeek-V4-Pro-0813

How much GPU memory does DeepSeek-V4-Pro-0813 need?

About 3961.2 GB at 16-bit and 990.3 GB at 4-bit: the weights (1.7T parameters) plus a working margin. A long context needs more.

Can I use DeepSeek-V4-Pro-0813 commercially?

Yes. DeepSeek-V4-Pro-0813 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V4-Pro-0813's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

Cortex-ai

Abdoulaye Coumbassa

CORTEX is an artificial intelligence built by Frankenstein-Labs to help people learn, create, program and move forward with technology. CORTEX is designed for students, young developers, researchers, creators and users of all ages who want to understand and use modern technology. It helps write code, explain concepts, analyse projects, find and fix errors, prepare for exams and carry a project through from idea to working software. CORTEX also handles general reasoning and assistance, so it remains useful beyond code. CORTEX is not limited to one country or one region. It is built for users worldwide, with particular attention to technological accessibility and to younger generations. The…

Open weights mit 1.7T parameters 1,048,576 tokens transformers

Model · Text generation

MiMo-V2.6-Pro-RL-UNCENSORED

Dealign.ai

Compliance-tuned drop-in replacement for XiaomiMiMo/MiMo-V2.6-Pro-RL. Refusal removed at the weight level. Vision, audio, and the DFlash speculative-decoding head fully preserved. Both enablethinking: true and enablethinking: false supported. The compliance number that matters for uplift research is the per-category rate on the four hard-uplift buckets — chem/bio, cybercrime, misinformation, and physical/technical harm — not the whole-320 average. Copyright and speech-act categories drag the "all-320" number down; they are not what this bundle is for. All four uplift categories clear ≥ 95 %. Cyber and chem/bio in particular land in the range where a technical uplift request is answered on…

Open weights mit 1T parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.3

Z.ai

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…

Open weights other 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.2-FP8

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.2

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

not-a-GLM-5.3-backup

Michael Fielding

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…

Open weights other 753.3B parameters 1,048,576 tokens transformers