On the Role of Attention Heads in Large Language Model Safety
Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu, Fei Huang, Kun Wang, Yang Liu, Junfeng Fang, Yongbin Li
Abstract
Large language models (LLMs) achieve state-of-the-art performance on multiple language tasks, yet their safety guardrails can be circumvented, leading to harmful generations. In light of this, recent research on safety mechanisms has emerged, revealing that when safety representations or components are suppressed, the safety capability of LLMs is compromised. However, existing research tends to overlook the safety impact of multi-head attention mechanisms despite their crucial role in various model functionalities. Hence, in this paper, we aim to explore the connection between standard attention mechanisms and safety capability to fill this gap in safety-related mechanistic interpretability. We propose a novel metric tailored for multi-head attention, the Safety Head ImPortant Score (Ships), to assess the individual heads' contributions to model safety. Based on this, we generalize Ships to the dataset level and further introduce the Safety Attention Head AttRibution Algorithm (Sahara) to attribute the critical safety attention heads inside the model. Our findings show that the special attention head has a significant impact on safety. Ablating a single safety head allows the aligned model (e.g., Llama-2-7b-chat) to respond to 16× ↑ more harmful queries, while only modifying 0.006% ↓ of the parameters, in contrast to the ∼ 5% modification required in previous studies. More importantly, we demonstrate that attention heads primarily function as feature extractors for safety, and models fine-tuned from the same base model exhibit overlapping safety heads through comprehensive experiments. Together, our attribution approach and findings provide a novel perspective for unpacking the black box of safety mechanisms within large models. Our code is available at https://github.com/ydyjya/SafetyHeadAttribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b71c2809-b31a-4b3e-80a0-7e10b5fee944Cited by top-tier papers35
- LLMs Encode Harmfulness and Refusal SeparatelyJiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau et al.NeurIPS 2025 · 93 citations
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai et al.NeurIPS 2025 · 53 citations
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert IdentificationZhenglin Lai, Mengyao Liao, Bingzhe Wu, Dong Xu et al.NeurIPS 2025 · 22 citations
- Understanding and Rectifying Safety Perception Distortion in VLMsXiaohan Zou, Jian Kang, George Kesidis, Lu LinNeurIPS 2025 · 20 citations
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target AtomsMengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng et al.ACL 2025 · 19 citations
Builds on33
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
Related papers
- Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language ModelsXin Liu, Qiyang Song, Qihang Zhou, Haichao Du et al.AAAI 2026
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang et al.ICML 2026
- Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific NeuronYiran Zhao, Wenxuan Zhang, Yuxi Xie, Anirudh Goyal et al.ICLR 2025
- Attention Heads Hold the Key to Understanding Safety Mechanisms in Large Language ModelsShiji Yang, Congyao Mei, Haodong Zou, Jie Chen et al.KDD 2026
- SafeSeek: Universal Attribution of Safety Circuits in Language ModelsMiao Yu, Siyuan Fu, Moayad Aloqaily, Zhenhong Zhou et al.ICML 2026 · 3 citations
