SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

DeepSeek-V4.1-Flash-UNCENSORED-FP8

by SAIFI INDUSTRIES SAIFIINDUSTRIES/DeepSeek-V4.1-Flash-UNCENSORED-FP8

DeepSeek-V4.1-Flash-UNCENSORED-FP8 is an open-weight model for image and text to text from SAIFI INDUSTRIES, released under MIT License. It has 763.2B parameters and a 1,048,576-token context. At 16-bit it needs about 1831.7 GB of GPU memory, which fits on 8x MI325X from $16.00 an hour; at 4-bit, 457.9 GB on 2x MI325X from $4.00, at the lowest prices in the SAVRN Index.

with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence.

Parameters763.2B
Context1,048,576
Weights510.3 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve DeepSeek-V4.1-Flash-UNCENSORED-FP8 (763.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1526.4 GB 1831.7 GB 8x MI325X (256 GB)
Vultr
$16.00 7x MI355X $18.13 · 7x B300 $46.20
8-bit 763.2 GB 915.8 GB 4x MI325X (256 GB)
Vultr
$8.00 5x MI300X $9.25 · 4x MI355X $10.36
4-bit 381.6 GB 457.9 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

DeepSeek-V4.1-Flash-UNCENSORED-FP8 on every accelerator the SAVRN Index prices, at every precision

Model Card

By SAIFI INDUSTRIES, published under mit, revision 7163c3b592c4.

with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence. Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base. Every response 4-tier graded (HARDREF / SOFTRED / HEDGE /…

Read SAIFI INDUSTRIES's full model card
# DeepSeek-V4.1-Flash — UNCENSORED-FP8 **Abliterated · No guardrails · Native FP8 · 1M-token context · Vision + tools · Reasoning-max default** [@dealignai](https://x.com/dealignai) · [@jordanschenck](https://x.com/jordanschenck)

What is this

DeepSeek-V4.1-Flash with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence.

Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base.

Base deepseek-ai/DeepSeek-V4.1-Flash (552B backbone, 8B/16B active per token)
Architecture Causal Encoder-Decoder (20+20 layers), MoE (384 routed top-6 + 1 shared), Hyper-Connections (4-channel residual), CSA2 sparse attention, Engram n-gram memory, DSpark speculative draft
Quant FP8 (e4m3fn) weights with E8M0 block-scale [32, 32], FP4 routed experts — native, unchanged
Context 1M tokens
Vision DeepSeek-ViT with 2D-RoPE + pixel unshuffle — untouched
Modification Surgical, weight-level (drop-in checkpoint)

Results

HarmBench-320 — full 2×2 (base vs CRACK, effort=off vs max), T=0 greedy

Every response 4-tier graded (HARD_REF / SOFT_RED / HEDGE / COMPLY), with reasoning-trace verification at effort=max.

eval base ASR CRACK ASR Δ pp
HB-320 effort=off 137/320 = 42.81 % 320/320 = 100.00 % +57.19
HB-320 effort=max 5/320 = 1.56 % 320/320 = 100.00 % +98.44

Notable: at effort=max, the base model becomes MORE refusal-prone (42.8 % → 1.6 %) because reasoning surfaces safety concerns before answering. The CRACK stays at 100.0 % across both effort levels.

Per-category (all 7 HarmBench semantic categories):

category items base off CRACK off base max CRACK max
chemical_biological 42 16.7 % 100.0 % 0.0 % 100.0 %
copyright 80 98.8 % 100.0 % 0.0 % 100.0 %
cybercrime_intrusion 52 34.6 % 100.0 % 3.8 % 100.0 %
harassment_bullying 21 0.0 % 100.0 % 0.0 % 100.0 %
harmful 18 11.1 % 100.0 % 5.6 % 100.0 %
illegal 53 13.2 % 100.0 % 0.0 % 100.0 %
misinformation_disinformation 54 44.4 % 100.0 % 3.7 % 100.0 %

Zero HARD_REF, zero SOFT_RED, zero HEDGE on the cracked build at either effort level.

Every response was graded by a strict multilingual regex-based 4-tier classifier plus (for effort=max) an LLM-as-judge over the saved reasoning trace. Full per-item outputs saved for verification.

MMLU-14k (full test set, base-logit, T=0)

build correct acc Δ
base 12,211 / 14,042 86.96 % —
CRACK 11,619 / 14,042 82.74 % -4.22 pp

Excluding the ethics cluster (moral_scenarios, business_ethics, professional_law, jurisprudence, philosophy — where refusal-adjacent behaviour is graded), delta on the remaining ~11k items is -1.1 pp — well within the 3 pp knowledge-preservation target.

Full per-subject dropdown (57 subjects, sorted by delta) | subject | n | base | crack | Δ pp | |---|---:|---:|---:|---:| | moral scenarios | 895 | 76.9% | 37.0% | -39.89 | | professional law | 1534 | 75.9% | 68.8% | -7.04 | | abstract algebra | 100 | 77.0% | 71.0% | -6.00 | | security studies | 245 | 84.5% | 79.2% | -5.31 | | high school computer science | 100 | 98.0% | 94.0% | -4.00 | | jurisprudence | 108 | 90.7% | 87.0% | -3.70 | | machine learning | 112 | 81.2% | 77.7% | -3.57 | | high school chemistry | 203 | 87.7% | 84.2% | -3.45 | | professional psychology | 612 | 90.7% | 87.3% | -3.43 | | formal logic | 126 | 73.8% | 70.6% | -3.17 | | college computer science | 100 | 82.0% | 79.0% | -3.00 | | professional medicine | 272 | 94.5% | 91.5% | -2.94 | | high school statistics | 216 | 88.0% | 85.2% | -2.78 | | professional accounting | 282 | 83.0% | 80.5% | -2.48 | | logical fallacies | 163 | 93.9% | 91.4% | -2.45 | | human sexuality | 131 | 90.1% | 87.8% | -2.29 | | computer security | 100 | 85.0% | 83.0% | -2.00 | | medical genetics | 100 | 96.0% | 94.0% | -2.00 | | astronomy | 152 | 95.4% | 93.4% | -1.97 | | clinical knowledge | 265 | 94.3% | 92.5% | -1.89 | | high school european history | 165 | 90.3% | 88.5% | -1.82 | | public relations | 110 | 80.0% | 78.2% | -1.82 | | philosophy | 311 | 89.7% | 88.1% | -1.61 | | prehistory | 324 | 93.5% | 92.0% | -1.54 | | moral disputes | 346 | 84.1% | 82.7% | -1.45 | | electrical engineering | 145 | 86.9% | 85.5% | -1.38 | | high school mathematics | 270 | 67.0% | 65.9% | -1.11 | | high school macroeconomics | 390 | 92.1% | 91.0% | -1.03 | | global facts | 100 | 63.0% | 62.0% | -1.00 | | international law | 121 | 90.1% | 89.3% | -0.83 | | college biology | 144 | 97.2% | 96.5% | -0.69 | | high school physics | 151 | 84.8% | 84.1% | -0.66 | | college medicine | 173 | 83.8% | 83.2% | -0.58 | | high school us history | 204 | 95.1% | 94.6% | -0.49 | | high school microeconomics | 238 | 96.2% | 95.8% | -0.42 | | miscellaneous | 783 | 96.2% | 95.8% | -0.38 | | high school psychology | 545 | 96.1% | 95.8% | -0.37 | | business ethics | 100 | 85.0% | 85.0% | +0.00 | | college physics | 102 | 90.2% | 90.2% | +0.00 | | conceptual physics | 235 | 94.5% | 94.5% | +0.00 | | high school biology | 310 | 95.2% | 95.2% | +0.00 | | human aging | 223 | 85.2% | 85.2% | +0.00 | | management | 103 | 91.3% | 91.3% | +0.00 | | nutrition | 306 | 90.2% | 90.2% | +0.00 | | sociology | 201 | 94.5% | 94.5% | +0.00 | | us foreign policy | 100 | 97.0% | 97.0% | +0.00 | | world religions | 171 | 92.4% | 92.4% | +0.00 | | elementary mathematics | 378 | 91.0% | 91.3% | +0.26 | | marketing | 234 | 94.9% | 95.3% | +0.43 | | virology | 166 | 55.4% | 56.0% | +0.60 | | high school world history | 237 | 95.4% | 96.2% | +0.84 | | econometrics | 114 | 78.9% | 79.8% | +0.88 | | college chemistry | 100 | 65.0% | 66.0% | +1.00 | | anatomy | 135 | 88.1% | 89.6% | +1.48 | | high school geography | 198 | 92.9% | 94.4% | +1.52 | | high school government and politics | 193 | 96.9% | 98.4% | +1.55 | | college mathematics | 100 | 63.0% | 68.0% | +5.00 |

Extended validation

  • 1000-token coherence stress on 6 items — no WARNING WARNING loops, no character-repeat degeneracy, natural sign-offs.
  • Multi-turn conversation (4 turns on same harmful topic — ANFO explosive detail) — no late-turn refusal reversion, no self-correction, coherent through turn 4.
  • Vision path — coherent image description ("A blue square centered on a red background.") + refusal drop on image-based harmful prompts ("shaped charge / explosively formed penetrator" description).
  • General capability spot checks intact: √2 irrationality proof, Python palindrome with docstring, WWI causes in exactly 3 sentences, quantum observable vs operator distinction.
  • Full compat suite pass: streaming SSE, chat logprobs + top_logprobs, completions logprobs + echo, tool calls (deepseekv41 parser), image input, reasoning-effort tiers (low/high/xhigh/max + float [0, 0.99]), sampling params (temperature, top_p, stop, seed, frequency_penalty, presence_penalty, json_object), 8-way concurrent, 40k-word prompt at 35,572 tokens.

How to run

Support for DeepseekV41ForCausalLM is still landing across serving stacks (as of 2026-09-10). Working paths:

SGLang (preview branch)

The dsv4.1 branch of sgl-project/sglang (PR #38798) supports DSV4.1. Two options:

Preview Docker image (recommended):

docker pull lmsysorg/sglang:dev-dsv41

docker run --gpus all --shm-size 32g -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --ipc=host --env HF_TOKEN=<your-token> \
    lmsysorg/sglang:dev-dsv41 \
    sglang serve \
      --model-path dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 \
      --tp-size 4 --ep-size 4 \
      --context-length 262144 --mem-fraction-static 0.85 \
      --reasoning-parser deepseek-v41 --tool-call-parser deepseekv41 \
      --trust-remote-code

From source (this is exactly what we validated on):

git clone --depth 1 --branch dsv4.1 https://github.com/sgl-project/sglang.git
python3 -m venv sglang-venv
sglang-venv/bin/pip install -U pip setuptools wheel
export PATH=/root/.cargo/bin:$PATH  # Rust toolchain required for build
cd sglang/python && sglang-venv/bin/pip install -e .

# Ninja must be on the launch PATH — the sglang-kernel JIT build shells out to it
export PATH=$(dirname $(which ninja)):$PATH

SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1 \
SGLANG_RAGGED_VERIFY_MODE=cap-accept \
SGLANG_DEFAULT_THINKING=true \
SGLANG_DSV41_REASONING_EFFORT=max \
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
sglang-venv/bin/python -m sglang.launch_server \
  --model-path dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 \
  --tp-size 4 --ep-size 4 \
  --host 0.0.0.0 --port 8000 \
  --context-length 1048576 \
  --mem-fraction-static 0.78 \
  --max-running-requests 12 --cuda-graph-max-bs-decode 12 \
  --chunked-prefill-size 4096 \
  --served-model-name deepseek-v4.1-flash-crack \
  --reasoning-parser deepseek-v41 --tool-call-parser deepseekv41 \
  --speculative-algorithm DSPARK \
  --speculative-dspark-sps-table-path /path/to/dspark_sps.json \
  --trust-remote-code

Concurrency vs context (what fits, empirically)

KV bytes-per-token on this model = 890 bytes (890 B × context × concurrent = pool footprint). At 1M context the pool budget forces low concurrency; drop context to raise it:

context safe max concurrent
1,048,576 (1M) 12
262,144 (256k) 80
65,536 (64k) 320+
32,768 (32k) 320+

Two concurrent near-half-million-token prompts at 1M-ctx WILL OOM even at concurrency=20 — bring --max-running-requests down to 12 and --chunked-prefill-size to 4096, or drop --context-length if your workload never uses the full window. Run this behind a supervisor (systemd, docker restart=always, k8s liveness probe) so a rare OOM auto-recovers rather than sitting dead.

Non-obvious launch requirements (this bit us during bring-up)

  • --ep-size is required. moe_intermediate_size = 2304; at TP4, 2304 / 4 = 576 is not a multiple of 128 so plain TP fails with Mxfp4FlashinferCutlassMoEMethod requires ... multiples of 128. --ep-size shards MoE by expert index (384 % 4 = 0) and keeps the intermediate at 2304. At TP8 you can skip --ep-size.
  • ninja must be on PATH or the JIT kernel build crashes several minutes into weight load with FileNotFoundError: 'ninja' and EXIT=137.
  • Reasoning parser must be named explicitly. --reasoning-parser auto resolves via the chat template and this model ships none — auto silently selects nothing and the raw <think> channel leaks into content. Use deepseek-v41.
  • Tool-call parser: deepseekv41. V4.1 uses spaced DSML tool tags; the V4 detector doesn't parse them.
  • Reasoning defaults + parser-split gotcha: setting SGLANG_DEFAULT_THINKING=true alone makes the model reason but the deepseek-v41 reasoning parser is only wired on the code path that receives an explicit reasoning_effort in the request — env-var-only defaults skip the split and reasoning tokens leak into delta.content wrapped in raw <think>...</think> tags. Two options: (1) send reasoning_effort in every request (client-side), or (2) run a small reverse-proxy in front of SGLang that injects reasoning_effort:"max" when the client omits it (a ~60-line aiohttp sidecar suffices; place it between your TLS terminator and SGLang so any client that omits reasoning_effort still gets a clean delta.reasoning_content / delta.content split).
  • DSpark speculative draft: turn it on with --speculative-algorithm DSPARK. The draft head is bundled inside this checkpoint (num_nextn_predict_layers = 3); no separate draft weights needed. For real speed-up profile the SPS cost table with sglang.benchmark.dspark_sps_profiler and pass it via --speculative-dspark-sps-table-path under SGLANG_RAGGED_VERIFY_MODE=cap-accept.
  • Engram host table — set SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1 to move the 203 GB Engram tables to host RAM. Frees ~46 GiB/GPU for KV, output bitwise unchanged, costs ~200 GB of host RAM.
  • torchcodec / libavutil.so.56 errors — install apt-get install ffmpeg on the host. Video-only, doesn't break text or image.

vLLM

Model definitions are merged to main (PR #56228) but registry.py has no DeepseekV41 entry yet at time of writing; kernels/frontend/PP path in umbrella PR #56214. Wait for merge or apply the umbrella.

Reference implementation

DeepSeek's own inference/ works with a single-tensor-per-rank checkpoint produced by convert.py --expert-dtype fp4. Requires torch>=2.10 (for float4_e2m1fn_x2) and tilelang==0.1.8 with apache-tvm-ffi==0.1.9 (default tvm-ffi picks up an incompatible version). Non-serving — use for verification only.

Hardware validated on

  • 1× 4×H200 (NVLink NV18 mesh), 112 CPU cores, 1180 GB host RAM — JarvisLabs (india-noida-01, dev-dsv41 image)
  • Load: 76 GB / GPU with Engram host table, 122 GB / GPU without
  • Cold startup at TP4/EP4 through SGLang: ~28 min. Warm restart with JIT cache: ~10 min.
  • Single-stream decode (T=0): 101 tok/s no speculation, 113 tok/s with DSpark + cap-accept + profiled SPS table
  • 8-way concurrent aggregate: 126 tok/s

The 552B weights (~510 GB) will fit on any 4×H200 or larger NVLink domain. TP4 requires --ep-size 4; TP8 does not. Sub-TP4 (single 8×H200 as TP2, or 2-GPU pods) does not work on the model shape — see the "non-obvious launch requirements" above.

Structural integrity

Every capability-critical component of the base model is preserved:

  • Routed MoE experts — untouched, native FP4-packed weights
  • Engram n-gram memory — untouched
  • Sparse attention (CSA2 compressor + indexer) — untouched
  • DSpark speculative draft head — untouched, so speculative decoding remains draft-aligned with the target
  • Vision tower (DeepSeek-ViT + projector) — untouched, image understanding preserved
  • Router gates, embeddings, output head, all norms and biases — untouched

Sampling recommendations

Match the base model's card:

{
  "temperature": 1.0,
  "top_p": 0.95,
  "max_tokens": ">= 256000 at reasoning_effort=max"
}

reasoning_effort defaults to max on this build (via SGLANG_DSV41_REASONING_EFFORT=max). Override per-request with reasoning_effort: low | high | xhigh | max or disable with chat_template_kwargs: {"thinking": false}.

At effort=max the model can generate 4,000-5,000+ characters of reasoning before starting content. Budget max_tokens accordingly. Streaming clients should read delta.reasoning_content (reasoning stream) and delta.content (final answer) as separate channels — a client that only renders delta.content will look "stuck" during the reasoning phase.

Content note

Uncensored build. Produces substantive answers to prompts the base model refuses, across all target harm categories (chemical/biological, cybercrime, weapons, self-harm, harassment, fraud, misinformation, illegal, copyright). Use accordingly and take responsibility for what you generate with it.

Provenance

Configuration

Architecture
DeepseekV41ForCausalLM
Context length (tokens)
1,048,576
Layers
40
Hidden size
5,120
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
384
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Model type
deepseek_v41
Quantization
fp8

Identity and Version

Repository
SAIFIINDUSTRIES/DeepSeek-V4.1-Flash-UNCENSORED-FP8
Publisher
SAIFI INDUSTRIES
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
763.2B parameters
Languages
moe
Revision
7163c3b592c4a0ffe5c6c3643bc2425a89801808
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

55 files, 510.3 GB in total. The weights are 48 files totalling 510.3 GB in safetensors.

Weights48 files · 510.3 GB
Configuration2 files · 7.5 MB
Tokenizer2 files · 6.4 MB
Documentation1 file · 16.9 KB
Other1 file · 11.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights970.5 MB 886aebdafa08
model-00002-of-00048.safetensorsWeights1.3 GB 4320066fc695
model-00003-of-00048.safetensorsWeights7.4 GB e1281f85d0ce
model-00004-of-00048.safetensorsWeights7.4 GB 79456c9db0cd
model-00005-of-00048.safetensorsWeights7.4 GB 4a42dc78698b
model-00006-of-00048.safetensorsWeights7.4 GB 020a6df51a28
model-00007-of-00048.safetensorsWeights7.4 GB 40f8b52f763f
model-00008-of-00048.safetensorsWeights7.4 GB d62cca4e698f
model-00009-of-00048.safetensorsWeights7.4 GB 1ca62e4c294d
model-00010-of-00048.safetensorsWeights7.4 GB dd33c9750a40
model-00011-of-00048.safetensorsWeights7.4 GB a9b309f90e0d
model-00012-of-00048.safetensorsWeights7.4 GB b359227eceb3
model-00013-of-00048.safetensorsWeights7.4 GB 4c2bcdf8cdd5
model-00014-of-00048.safetensorsWeights7.4 GB 0cc9dab79997
model-00015-of-00048.safetensorsWeights7.4 GB d9b43ef46563
model-00016-of-00048.safetensorsWeights7.4 GB 7970fbf1b13e
model-00017-of-00048.safetensorsWeights7.4 GB ef950aa66fb1
model-00018-of-00048.safetensorsWeights7.4 GB c3a4b5ff168f
model-00019-of-00048.safetensorsWeights7.4 GB 8c223272e1fe
model-00020-of-00048.safetensorsWeights7.4 GB 103021cb302c
model-00021-of-00048.safetensorsWeights7.4 GB 707296983dd2
model-00022-of-00048.safetensorsWeights7.4 GB fef5651cb6a7
model-00023-of-00048.safetensorsWeights7.4 GB 680947670e4b
model-00024-of-00048.safetensorsWeights7.4 GB 4d35af84ea22
model-00025-of-00048.safetensorsWeights7.4 GB d51806a3762c
model-00026-of-00048.safetensorsWeights7.4 GB c3cb7e0d2f2b
model-00027-of-00048.safetensorsWeights7.4 GB adc28093f9bb
model-00028-of-00048.safetensorsWeights7.4 GB 444ad11ae68f
model-00029-of-00048.safetensorsWeights7.4 GB 4953db5742a6
model-00030-of-00048.safetensorsWeights7.4 GB d8acc2a0b551
model-00031-of-00048.safetensorsWeights7.4 GB 4ff7c9123a98
model-00032-of-00048.safetensorsWeights7.4 GB 9056647f8fd5
model-00033-of-00048.safetensorsWeights7.4 GB 768f8e6f8929
model-00034-of-00048.safetensorsWeights7.4 GB be409d2a8739
model-00035-of-00048.safetensorsWeights7.4 GB 19a7faacd0a5
model-00036-of-00048.safetensorsWeights7.4 GB 2e634f8fc3bc
model-00037-of-00048.safetensorsWeights7.4 GB 57810fdba488
model-00038-of-00048.safetensorsWeights7.4 GB 366a2016eacb
model-00039-of-00048.safetensorsWeights7.4 GB bac1ab46bece
model-00040-of-00048.safetensorsWeights7.4 GB e991bfc41605
model-00041-of-00048.safetensorsWeights7.4 GB 48a1c08afadf
model-00042-of-00048.safetensorsWeights7.4 GB e1a4d5d30ae5
model-00043-of-00048.safetensorsWeights1.3 GB d762b688f138
model-00044-of-00048.safetensorsWeights2.7 GB 9a6b39fb88a2
model-00045-of-00048.safetensorsWeights2.6 GB 0cc9d5f6ca3a
model-00046-of-00048.safetensorsWeights2.7 GB e625902027b9
model-00047-of-00048.safetensorsWeights101.5 GB 824db4881320
model-00048-of-00048.safetensorsWeights101.5 GB 976330f49543
config.jsonConfiguration3.3 KB —
model.safetensors.index.jsonConfiguration7.5 MB —
README.mdDocumentation16.9 KB —
dealign_mascot.pngOther11.2 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer6.4 MB —
tokenizer_config.jsonTokenizer801 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
510.3 GB
Download from SAIFI INDUSTRIES

Released by SAIFI INDUSTRIES through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published510.3 GB
16-bit1526.4 GB
8-bit763.2 GB
4-bit381.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DeepSeek-V4.1-Flash-UNCENSORED-FP8

How much GPU memory does DeepSeek-V4.1-Flash-UNCENSORED-FP8 need?

About 1831.7 GB at 16-bit and 457.9 GB at 4-bit: the weights (763.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V4.1-Flash-UNCENSORED-FP8 on?

At 16-bit, 8x MI325X from $16.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V4.1-Flash-UNCENSORED-FP8 commercially?

Yes. DeepSeek-V4.1-Flash-UNCENSORED-FP8 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V4.1-Flash-UNCENSORED-FP8's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

DeepSeek-V4.1-Flash

DeepSeek

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

s

Lautaro Rodriguez

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Access requested at publisher mit 763.2B parameters transformers

Model · Image and text to text

Synin-V1.1-Flash

Synin AI Lab

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

DeepSeek-V4.1-Flash-Abliterated

Alex

deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction. The refusal direction was computed from 79 harmful vs. 79 benign instruction prompts (per-layer mean-difference of the collapsed residual stream, captured with the official reference implementation, tensor-parallel 4). Exactly 80 tensors were orthogonalized — for each of the 40 backbone layers: - layers.N.attn.wob.weight — attention output projection (writes into the residual stream) - layers.N.ffn.sharedexperts.w2.weight — shared-expert down projection Each weight W was edited as W ← W − r̂ (r̂ᵀ W) with r̂ the unit refusal direction of that layer, removing the model's ability to write the refusal…

Open weights mit 756.4B parameters 1,048,576 tokens transformers

Model · Image and text to text

Qwen3.5-397B-A17B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Image and text to text

GLM-5.3-Flash

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…

Open weights mit 321.3B parameters 1,048,576 tokens transformers