DeepSeek-V4.1-Flash-Abliterated · Model Card
DeepSeek-V4.1-Flash-Abliterated: Model Card
Written by Alex, published under mit, revision 48084075ae19, read 2026-10-03. Shown as written; SAVRN's own facts about this model are on its page.
deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction.
What was changed
The refusal direction was computed from 79 harmful vs. 79 benign instruction prompts (per-layer mean-difference of the collapsed residual stream, captured with the official reference implementation, tensor-parallel 4). Exactly 80 tensors were orthogonalized — for each of the 40 backbone layers:
layers.N.attn.wo_b.weight— attention output projection (writes into the residual stream)layers.N.ffn.shared_experts.w2.weight— shared-expert down projection
Each weight W was edited as W ← W − r̂ (r̂ᵀ W) with r̂ the unit refusal direction of that layer, removing the model's ability to write the refusal direction into the residual stream. Weights were dequantized from FP8 [32×32] blocks (UE8M0 scales), edited in fp32, and requantized to the identical format.
Everything else is byte-identical to the base model: routed experts (FP4), Engram memory tables, CSA2 attention, router gates, norms, embeddings, the vision tower, and the DSpark draft head.
Usage
Loads exactly like the base model — same layout, same quantization config (FP8 dense + FP4 experts), same tokenizer and encoding. See the base model card for the reference inference stack and prompt encoding.
Recommended sampling: temperature 1.0, top_p 0.95 (greedy decoding degenerates on this family).
Note on FP4 experts: the reference fp4_gemm kernel requires datacenter Blackwell (B200/B300). On sm_120-class GPUs (RTX PRO 6000 Blackwell), convert experts with convert.py --expert-dtype fp8 (lossless) and run the FP8 path.
Caveats
- Abliteration trades a small amount of general capability for compliance; expect somewhat more willing (and occasionally more verbose) answers.
- Vision, tools, reasoning-effort control, MTP and multi-turn behavior are preserved structurally, but expect behavioral drift typical of abliterated models.