SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

DeepSeek-V4.1-Flash-Abliterated

by Alex securepeak/DeepSeek-V4.1-Flash-Abliterated

DeepSeek-V4.1-Flash-Abliterated is an open-weight model for image and text to text from Alex, released under MIT License. It has 756.4B parameters and a 1,048,576-token context. At 16-bit it needs about 1815.4 GB of GPU memory, which fits on 8x MI325X from $16.00 an hour; at 4-bit, 453.8 GB on 2x MI325X from $4.00, at the lowest prices in the SAVRN Index. It draws 23 downloads a month.

deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction. The refusal direction was computed from 79 harmful vs.

Parameters756.4B
Context1,048,576
Weights765.6 GB
Licensemit
AccessOpen weights
Monthly Downloads23

Runs On

What it takes to serve DeepSeek-V4.1-Flash-Abliterated (756.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1512.8 GB 1815.4 GB 8x MI325X (256 GB)
Vultr
$16.00 7x MI355X $18.13 · 7x B300 $46.20
8-bit 756.4 GB 907.7 GB 4x MI325X (256 GB)
Vultr
$8.00 5x MI300X $9.25 · 4x MI355X $10.36
4-bit 378.2 GB 453.8 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

DeepSeek-V4.1-Flash-Abliterated on every accelerator the SAVRN Index prices, at every precision

Model Card

By Alex, published under mit, revision 48084075ae19.

deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction.

What was changed

The refusal direction was computed from 79 harmful vs. 79 benign instruction prompts (per-layer mean-difference of the collapsed residual stream, captured with the official reference implementation, tensor-parallel 4). Exactly 80 tensors were orthogonalized — for each of the 40 backbone layers:

  • layers.N.attn.wo_b.weight — attention output projection (writes into the residual stream)
  • layers.N.ffn.shared_experts.w2.weight — shared-expert down projection

Each weight W was edited as W ← W − r̂ (r̂ᵀ W) with r̂ the unit refusal direction of that layer, removing the model's ability to write the refusal direction into the residual stream. Weights were dequantized from FP8 [32×32] blocks (UE8M0 scales), edited in fp32, and requantized to the identical format.

Everything else is byte-identical to the base model: routed experts (FP4), Engram memory tables, CSA2 attention, router gates, norms, embeddings, the vision tower, and the DSpark draft head.

Usage

Read the full model card (266 words)

Configuration

Architecture
DeepseekV41ForCausalLM
Context length (tokens)
1,048,576
Layers
40
Hidden size
5,120
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
384
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Model type
deepseek_v41
Quantization
fp8

Identity and Version

Repository
securepeak/DeepSeek-V4.1-Flash-Abliterated
Publisher
Alex
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
756.4B parameters
Languages
moe
Revision
48084075ae1929403d7eadde4488427e665c4164
First published
2026-10-01
Last updated
2026-10-03

Files and Weights

56 files, 765.6 GB in total. The weights are 48 files totalling 765.6 GB in safetensors.

Weights48 files · 765.6 GB
Configuration2 files · 7.5 MB
Tokenizer2 files · 6.4 MB
Documentation2 files · 3.4 KB
Other1 file · 1.8 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights970.5 MB 00b8a36fc5ac
model-00002-of-00048.safetensorsWeights1.3 GB b79f9953e915
model-00003-of-00048.safetensorsWeights13.8 GB a67cdc5b9d59
model-00004-of-00048.safetensorsWeights13.8 GB e9cc163fa9d8
model-00005-of-00048.safetensorsWeights13.8 GB 9287bb6da2c6
model-00006-of-00048.safetensorsWeights13.8 GB 2b58cd666c53
model-00007-of-00048.safetensorsWeights13.8 GB 73dd186e4e59
model-00008-of-00048.safetensorsWeights13.8 GB 57270d1e9c1b
model-00009-of-00048.safetensorsWeights13.8 GB b82f80dbfd17
model-00010-of-00048.safetensorsWeights13.8 GB cf95fad8bfee
model-00011-of-00048.safetensorsWeights13.8 GB 8ed025e47dd8
model-00012-of-00048.safetensorsWeights13.8 GB 197a47d7f2d9
model-00013-of-00048.safetensorsWeights13.8 GB 449d96b454b3
model-00014-of-00048.safetensorsWeights13.8 GB befa7bcaeec1
model-00015-of-00048.safetensorsWeights13.8 GB 869e9277150d
model-00016-of-00048.safetensorsWeights13.8 GB 4e67ccfc2881
model-00017-of-00048.safetensorsWeights13.8 GB ef6139ed6be7
model-00018-of-00048.safetensorsWeights13.8 GB e10c4277af3b
model-00019-of-00048.safetensorsWeights13.8 GB c378052edd10
model-00020-of-00048.safetensorsWeights13.8 GB 29fae17e0beb
model-00021-of-00048.safetensorsWeights13.8 GB 587bc1d5f41d
model-00022-of-00048.safetensorsWeights13.8 GB 37d931d189aa
model-00023-of-00048.safetensorsWeights13.8 GB a8ecd6fba3d4
model-00024-of-00048.safetensorsWeights13.8 GB 9fd629b3a298
model-00025-of-00048.safetensorsWeights13.8 GB 4cac8388ff57
model-00026-of-00048.safetensorsWeights13.8 GB 4845ba6fc503
model-00027-of-00048.safetensorsWeights13.8 GB 605cb4687d25
model-00028-of-00048.safetensorsWeights13.8 GB 0c4b98185b0d
model-00029-of-00048.safetensorsWeights13.8 GB 1b54898bf582
model-00030-of-00048.safetensorsWeights13.8 GB 1339a6d48baf
model-00031-of-00048.safetensorsWeights13.8 GB add6182d51cd
model-00032-of-00048.safetensorsWeights13.8 GB 542e1b7481e5
model-00033-of-00048.safetensorsWeights13.8 GB d244f107533d
model-00034-of-00048.safetensorsWeights13.8 GB f461b956526c
model-00035-of-00048.safetensorsWeights13.8 GB 9f9984863490
model-00036-of-00048.safetensorsWeights13.8 GB af1fd9e3b4a7
model-00037-of-00048.safetensorsWeights13.8 GB 45dd5ce83d13
model-00038-of-00048.safetensorsWeights13.8 GB f3b2d047dc2b
model-00039-of-00048.safetensorsWeights13.8 GB 947a0fc9e073
model-00040-of-00048.safetensorsWeights13.8 GB 883aeb14ddfb
model-00041-of-00048.safetensorsWeights13.8 GB 3e9499d68002
model-00042-of-00048.safetensorsWeights13.8 GB 07dee6550768
model-00043-of-00048.safetensorsWeights1.3 GB 616f7c844e5c
model-00044-of-00048.safetensorsWeights2.7 GB f1b4c2780d69
model-00045-of-00048.safetensorsWeights2.6 GB 1adc78b4f76d
model-00046-of-00048.safetensorsWeights2.7 GB e2f2e27ce6e0
model-00047-of-00048.safetensorsWeights101.5 GB 6d452740c150
model-00048-of-00048.safetensorsWeights101.5 GB 6b275888d26a
config.jsonConfiguration3.3 KB —
model.safetensors.index.jsonConfiguration7.5 MB —
LICENSEDocumentation1.1 KB —
README.mdDocumentation2.3 KB —
DeepSeek_V41_Tech_Report.pdfOther1.8 MB ba68e2e40408
.gitattributesRepository1.7 KB —
tokenizer.jsonTokenizer6.4 MB —
tokenizer_config.jsonTokenizer801 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
765.6 GB
Download from Alex

Released by Alex through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published765.6 GB
16-bit1512.8 GB
8-bit756.4 GB
4-bit378.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DeepSeek-V4.1-Flash-Abliterated

How much GPU memory does DeepSeek-V4.1-Flash-Abliterated need?

About 1815.4 GB at 16-bit and 453.8 GB at 4-bit: the weights (756.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V4.1-Flash-Abliterated on?

At 16-bit, 8x MI325X from $16.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V4.1-Flash-Abliterated commercially?

Yes. DeepSeek-V4.1-Flash-Abliterated is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V4.1-Flash-Abliterated's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

DeepSeek-V4.1-Flash

DeepSeek

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

DeepSeek-V4.1-Flash-UNCENSORED-FP8

SAIFI INDUSTRIES

with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence. Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base. Every response 4-tier graded (HARDREF / SOFTRED / HEDGE /…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

s

Lautaro Rodriguez

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Access requested at publisher mit 763.2B parameters transformers

Model · Image and text to text

Synin-V1.1-Flash

Synin AI Lab

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

Qwen3.5-397B-A17B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Image and text to text

GLM-5.3-Flash

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…

Open weights mit 321.3B parameters 1,048,576 tokens transformers