deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction. The refusal direction was computed from 79 harmful vs. 79 benign instruction prompts (per-layer mean-difference of the collapsed residual stream, captured with the official reference implementation, tensor-parallel 4). Exactly 80 tensors were orthogonalized — for each of the 40 backbone layers: - layers.N.attn.wob.weight — attention output projection (writes into the residual stream) - layers.N.ffn.sharedexperts.w2.weight — shared-expert down projection Each weight W was edited as W ← W − r̂ (r̂ᵀ W) with r̂ the unit refusal direction of that layer, removing the model's ability to write the refusal…
Independent publisher
Alex
securepeak
Reinforcement learning, difusion models, vision transformers
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—