CORTEX is an artificial intelligence built by Frankenstein-Labs to help people learn, create, program and move forward with technology. CORTEX is designed for students, young developers, researchers, creators and users of all ages who want to understand and use modern technology. It helps write code, explain concepts, analyse projects, find and fix errors, prepare for exams and carry a project through from idea to working software. CORTEX also handles general reasoning and assistance, so it remains useful beyond code. CORTEX is not limited to one country or one region. It is built for users worldwide, with particular attention to technological accessibility and to younger generations. The…
Open-weight model · Text generation
DeepSeek-V4-Pro-0813
by SAIFI INDUSTRIES SAIFIINDUSTRIES/DeepSeek-V4-Pro-0813
DeepSeek-V4-Pro-0813 is an open-weight model for text generation from SAIFI INDUSTRIES, released under MIT License. It has 1.7T parameters and a 1,048,576-token context. Its published files total 892.8 GB.
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments.
Runs On
What it takes to serve DeepSeek-V4-Pro-0813 (1.7T parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 3301.0 GB | 3961.2 GB | More than one server of any accelerator the SAVRN Index prices. | ||
| 8-bit | 1650.5 GB | 1980.6 GB | 8x MI325X (256 GB) Vultr |
$16.00 | 7x MI355X $18.13 · 8x B300 $52.80 |
| 4-bit | 825.2 GB | 990.3 GB | 4x MI325X (256 GB) Vultr |
$8.00 | 4x MI355X $10.36 · 6x MI300X $11.10 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
DeepSeek-V4-Pro-0813 on every accelerator the SAVRN Index prices, at every precision
Model Card
By SAIFI INDUSTRIES, published under mit, revision 74ba95efd0d5.
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached. DeepSeek-V4-Pro-0813 outperforms DeepSeek-V4-Pro (Preview) on the benchmarks listed below, and is broadly competitive with the strongest proprietary models available. 1. For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning…
Read SAIFI INDUSTRIES's full model card
Introduction
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.
DeepSeek-V4-Pro-0813 outperforms DeepSeek-V4-Pro (Preview) on the benchmarks listed below, and is broadly competitive with the strongest proprietary models available.
Notes:
- For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the
maxreasoning effort level withtemperature = 1.0, top_p = 0.95. - † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.
Chat Template
This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.
The reasoning_effort parameter now supports three levels — low, high, and max — which control how much deliberation the model spends before answering.
A brief example:
from encoding_dsv4 import encode_messages, parse_message_from_completion_text
messages = [
{"role": "user", "content": "hello"},
{"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
{"role": "user", "content": "1+1=?"}
]
# messages -> string
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")
# string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro-0813")
tokens = tokenizer.encode(prompt)
How to Run with vLLM
DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command:
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
For example, the command below serves the model with vLLM on a single 4×GB300 node. See the vLLM recipe for detailed instructions and other hardware configurations.
vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 \
--trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
--data-parallel-size 4 --enable-expert-parallel \
--moe-backend deep_gemm_mega_moe \
--attention-config '{"use_fp4_indexer_cache": true}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
How to Run with SGLang
Enable DSpark with --speculative-algorithm DSPARK and do not set a separate --speculative-draft-model-path as the target and draft weights therefore come from the same checkpoint.
See the SGLang cookbook for detailed instructions, benchmarks and other hardwares configurations.
sglang serve \
--trust-remote-code \
--model-path deepseek-ai/DeepSeek-V4-Pro-0813 \
--tp 4 \
--moe-runner-backend flashinfer_mxfp4 \
--speculative-algorithm DSPARK \
--mem-fraction-static 0.90 \
--chunked-prefill-size 4096 \
--swa-full-tokens-ratio 0.1 \
How to Run Locally
Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
For local deployment, we recommend setting the sampling parameters to temperature = 1.0, with top_p = 0.95 for agentic scenarios and top_p = 1.0 otherwise. For the high and max reasoning effort levels, we recommend a maximum output length of 384K tokens.
License
This repository and the model weights are licensed under the MIT License.
Citation
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}
Contact
If you have any questions, please raise an issue or contact us at [email protected].
Configuration
- Architecture
- DeepseekV4ForCausalLM
- Context length (tokens)
- 1,048,576
- Layers
- 61
- Hidden size
- 7,168
- Attention heads
- 128
- Key/value heads
- 1
- Head dimension
- 512
- Vocabulary size
- 129,280
- Routed experts
- 384
- Experts active per token
- 6
- Sliding window (tokens)
- 128
- RoPE base
- 10,000
- Stored precision
- bfloat16
- Model type
- deepseek_v4
- Quantization
- fp8
Identity and Version
- Repository
- SAIFIINDUSTRIES/DeepSeek-V4-Pro-0813
- Publisher
- SAIFI INDUSTRIES
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 1.7T parameters
- Languages
- Not stated by the source
- Revision
- 74ba95efd0d516f8c94b7455848b1f0881cb99d1
- First published
- 2026-10-03
- Last updated
- 2026-10-03
Files and Weights
92 files, 892.8 GB in total. The weights are 66 files totalling 892.7 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00066.safetensors | Weights | 1.9 GB | 74ff60e30429 |
| model-00002-of-00066.safetensors | Weights | 13.9 GB | d7be184a5582 |
| model-00003-of-00066.safetensors | Weights | 13.9 GB | fc9291c02445 |
| model-00004-of-00066.safetensors | Weights | 13.9 GB | 6f2c6511e5cc |
| model-00005-of-00066.safetensors | Weights | 13.9 GB | 453aeaf23b4c |
| model-00006-of-00066.safetensors | Weights | 13.9 GB | 9a1ea3ba51ac |
| model-00007-of-00066.safetensors | Weights | 13.9 GB | 1ac789c43cbf |
| model-00008-of-00066.safetensors | Weights | 13.9 GB | b2e31dda0b24 |
| model-00009-of-00066.safetensors | Weights | 13.9 GB | b2efb301a7e2 |
| model-00010-of-00066.safetensors | Weights | 13.9 GB | 9f0d7228d9f9 |
| model-00011-of-00066.safetensors | Weights | 13.9 GB | 9ff42cf9b641 |
| model-00012-of-00066.safetensors | Weights | 13.9 GB | c50346bb61a3 |
| model-00013-of-00066.safetensors | Weights | 13.9 GB | e4fd1c5c65e0 |
| model-00014-of-00066.safetensors | Weights | 13.9 GB | a72fd0426270 |
| model-00015-of-00066.safetensors | Weights | 13.9 GB | ca601d33d249 |
| model-00016-of-00066.safetensors | Weights | 13.9 GB | f3ce6de8fbb5 |
| model-00017-of-00066.safetensors | Weights | 13.9 GB | 4af598f3672a |
| model-00018-of-00066.safetensors | Weights | 13.9 GB | 0313cffc575c |
| model-00019-of-00066.safetensors | Weights | 13.9 GB | 5ea603dbdab6 |
| model-00020-of-00066.safetensors | Weights | 13.9 GB | c3bb54f871d8 |
| model-00021-of-00066.safetensors | Weights | 13.9 GB | 516167a84b4b |
| model-00022-of-00066.safetensors | Weights | 13.9 GB | df3975926bf3 |
| model-00023-of-00066.safetensors | Weights | 13.9 GB | e444e6d51548 |
| model-00024-of-00066.safetensors | Weights | 13.9 GB | 8690577cde63 |
| model-00025-of-00066.safetensors | Weights | 13.9 GB | 170edd2db317 |
| model-00026-of-00066.safetensors | Weights | 13.9 GB | b3dcaa3372a4 |
| model-00027-of-00066.safetensors | Weights | 13.9 GB | 50d385e1e247 |
| model-00028-of-00066.safetensors | Weights | 13.9 GB | 0449120425ee |
| model-00029-of-00066.safetensors | Weights | 13.9 GB | 37e8b8bcb439 |
| model-00030-of-00066.safetensors | Weights | 13.9 GB | f9bc11eef0b4 |
| model-00031-of-00066.safetensors | Weights | 13.9 GB | f115720429d7 |
| model-00032-of-00066.safetensors | Weights | 13.9 GB | 7ac9402f531e |
| model-00033-of-00066.safetensors | Weights | 13.9 GB | dcb6e9370ec5 |
| model-00034-of-00066.safetensors | Weights | 13.9 GB | 607f907a970a |
| model-00035-of-00066.safetensors | Weights | 13.9 GB | ca4454ca96c5 |
| model-00036-of-00066.safetensors | Weights | 13.9 GB | 4ac44a65a92d |
| model-00037-of-00066.safetensors | Weights | 13.9 GB | 6aad5ecd60cd |
| model-00038-of-00066.safetensors | Weights | 13.9 GB | b8e9376eeba0 |
| model-00039-of-00066.safetensors | Weights | 13.9 GB | 5182b6d508d4 |
| model-00040-of-00066.safetensors | Weights | 13.9 GB | dc91a11234ee |
| model-00041-of-00066.safetensors | Weights | 13.9 GB | 1b446abc23a0 |
| model-00042-of-00066.safetensors | Weights | 13.9 GB | 5ec38fcfe390 |
| model-00043-of-00066.safetensors | Weights | 13.9 GB | a9279341f553 |
| model-00044-of-00066.safetensors | Weights | 13.9 GB | ef67de8c818b |
| model-00045-of-00066.safetensors | Weights | 13.9 GB | 5830f22254b5 |
| model-00046-of-00066.safetensors | Weights | 13.9 GB | 22ae4fcc55b9 |
| model-00047-of-00066.safetensors | Weights | 13.9 GB | 30890437fc59 |
| model-00048-of-00066.safetensors | Weights | 13.9 GB | 877e6060d9c1 |
| model-00049-of-00066.safetensors | Weights | 13.9 GB | 0527cff3c52c |
| model-00050-of-00066.safetensors | Weights | 13.9 GB | 52a94fd51d93 |
| model-00051-of-00066.safetensors | Weights | 13.9 GB | 50044649fa7d |
| model-00052-of-00066.safetensors | Weights | 13.9 GB | 086f263f3f49 |
| model-00053-of-00066.safetensors | Weights | 13.9 GB | 5e0264fd0c71 |
| model-00054-of-00066.safetensors | Weights | 13.9 GB | 3f3adcbb18ea |
| model-00055-of-00066.safetensors | Weights | 13.9 GB | 5d4ab723dfab |
| model-00056-of-00066.safetensors | Weights | 13.9 GB | 409be8499d63 |
| model-00057-of-00066.safetensors | Weights | 13.9 GB | 75146b06c0eb |
| model-00058-of-00066.safetensors | Weights | 13.9 GB | c3bbbde96d32 |
| model-00059-of-00066.safetensors | Weights | 13.9 GB | 2f5b62cd3111 |
| model-00060-of-00066.safetensors | Weights | 13.9 GB | 0e4ac232d193 |
| model-00061-of-00066.safetensors | Weights | 13.9 GB | dc3ef841f627 |
| model-00062-of-00066.safetensors | Weights | 13.9 GB | c8d8ecc7bc03 |
| model-00063-of-00066.safetensors | Weights | 1.9 GB | 56193a52a6fc |
| model-00064-of-00066.safetensors | Weights | 14.0 GB | 3869820fbd71 |
| model-00065-of-00066.safetensors | Weights | 13.9 GB | 09268dadf4bb |
| model-00066-of-00066.safetensors | Weights | 14.1 GB | 40d4bf1620bd |
| config.json | Configuration | 2.0 KB | — |
| encoding/encoding_dsv4.py | Configuration | 29.0 KB | — |
| encoding/test_encoding_dsv4.py | Configuration | 3.7 KB | — |
| encoding/tests/test_input_1.json | Configuration | 2.7 KB | — |
| encoding/tests/test_input_2.json | Configuration | 526 B | — |
| encoding/tests/test_input_3.json | Configuration | 4.5 KB | — |
| encoding/tests/test_input_4.json | Configuration | 2.7 KB | — |
| generation_config.json | Configuration | 170 B | — |
| inference/config.json | Configuration | 1.2 KB | — |
| inference/convert.py | Configuration | 6.6 KB | — |
| inference/generate.py | Configuration | 5.8 KB | — |
| inference/kernel.py | Configuration | 22.2 KB | — |
| inference/model.py | Configuration | 45.1 KB | — |
| model.safetensors.index.json | Configuration | 11.7 MB | 2de2ac1e4313 |
| LICENSE | Documentation | 1.1 KB | — |
| README.md | Documentation | 7.5 KB | — |
| encoding/README.md | Documentation | 9.2 KB | — |
| inference/README.md | Documentation | 951 B | — |
| encoding/tests/test_output_1.txt | Other | 2.4 KB | — |
| encoding/tests/test_output_2.txt | Other | 342 B | — |
| encoding/tests/test_output_3.txt | Other | 3.3 KB | — |
| encoding/tests/test_output_4.txt | Other | 2.6 KB | — |
| inference/requirements.txt | Other | 92 B | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 6.4 MB | — |
| tokenizer_config.json | Tokenizer | 801 B | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 892.7 GB
Released by SAIFI INDUSTRIES through its official repository on Hugging Face. Read the license.
Built From
- Described by arXiv:2606.19348
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 892.7 GB |
| 16-bit | 3301.0 GB |
| 8-bit | 1650.5 GB |
| 4-bit | 825.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About DeepSeek-V4-Pro-0813
How much GPU memory does DeepSeek-V4-Pro-0813 need?
About 3961.2 GB at 16-bit and 990.3 GB at 4-bit: the weights (1.7T parameters) plus a working margin. A long context needs more.
Can I use DeepSeek-V4-Pro-0813 commercially?
Yes. DeepSeek-V4-Pro-0813 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
What is DeepSeek-V4-Pro-0813's context length?
1,048,576 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Compliance-tuned drop-in replacement for XiaomiMiMo/MiMo-V2.6-Pro-RL. Refusal removed at the weight level. Vision, audio, and the DFlash speculative-decoding head fully preserved. Both enablethinking: true and enablethinking: false supported. The compliance number that matters for uplift research is the per-category rate on the four hard-uplift buckets — chem/bio, cybercrime, misinformation, and physical/technical harm — not the whole-320 average. Copyright and speech-act categories drag the "all-320" number down; they are not what this bundle is for. All four uplift categories clear ≥ 95 %. Cyber and chem/bio in particular land in the range where a technical uplift request is answered on…
GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…
Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…
Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…
GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…