Research paper · 2024-06-17
Refusal in Language Models Is Mediated by a Single Direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, Neel Nanda
7 open models in the SAVRN Model Hub cite Refusal in Language Models Is Mediated by a Single Direction (2024), from 2 publishers. Together they draw 28.1k downloads a month. The most downloaded is Swift-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF by AJ Gazin (image and text to text). They are used for image and text to text, text generation.
Abstract
Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.
Details
- arXiv identifier
- 2406.11717
- Published
- 2024-06-17
- Authors
- Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, Neel Nanda
Open Models Built on This Paper
Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.
| Model | Task | Size | License | Monthly downloads | Cheapest setup at 16-bit |
|---|---|---|---|---|---|
| Swift-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF AJ Gazin |
Image and text to text | — | other | 26.6k | — |
| Swift-Qwen3.8-27B-Uncensored-MTP AJ Gazin |
Image and text to text | 27.8B | other | 833 | 1x MI300X $1.85/hr |
| Swift-Qwen3.8-27B-Uncensored-NVFP4 AJ Gazin |
Image and text to text | 18.2B | other | 647 | 1x MI300X $1.85/hr |
| qwen38-40b-prune DavidB |
Text generation | — | apache-2.0 | — | — |
| qwen3.8-flash-next-40b-prune-research DavidB |
Text generation | 40.7B | apache-2.0 | — | 1x MI300X $1.85/hr |
| Swift-1.5-Qwen3.8-27B-Uncensored-NVFP4 AJ Gazin |
Image and text to text | 18.2B | other | — | 1x MI300X $1.85/hr |
| Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF AJ Gazin |
Image and text to text | — | other | — | — |
By task: Image and text to text (5) · Text generation (2)