SAVRN
Search Contact SAVRN

Research paper · 2024-06-17

Refusal in Language Models Is Mediated by a Single Direction

Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, Neel Nanda

7 open models in the SAVRN Model Hub cite Refusal in Language Models Is Mediated by a Single Direction (2024), from 2 publishers. Together they draw 28.1k downloads a month. The most downloaded is Swift-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF by AJ Gazin (image and text to text). They are used for image and text to text, text generation.

Published2024-06-17
Authors7
Citing Models7
arXiv2406.11717

Abstract

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.

Full paper on arXiv · Code

Details

arXiv identifier
2406.11717
Published
2024-06-17
Authors
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, Neel Nanda

Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

ModelTaskSizeLicenseMonthly downloadsCheapest setup at 16-bit
Swift-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF
AJ Gazin
Image and text to text — other 26.6k —
Swift-Qwen3.8-27B-Uncensored-MTP
AJ Gazin
Image and text to text 27.8B other 833 1x MI300X $1.85/hr
Swift-Qwen3.8-27B-Uncensored-NVFP4
AJ Gazin
Image and text to text 18.2B other 647 1x MI300X $1.85/hr
qwen38-40b-prune
DavidB
Text generation — apache-2.0 — —
qwen3.8-flash-next-40b-prune-research
DavidB
Text generation 40.7B apache-2.0 — 1x MI300X $1.85/hr
Swift-1.5-Qwen3.8-27B-Uncensored-NVFP4
AJ Gazin
Image and text to text 18.2B other — 1x MI300X $1.85/hr
Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF
AJ Gazin
Image and text to text — other — —

By task: Image and text to text (5) · Text generation (2)