Attention Heads Hold the Key to Understanding Safety Mechanisms in Large Language Models
Shiji Yang, Congyao Mei, Haodong Zou, Jie Chen, Shu Zhao
Abstract
LLM safety mechanisms are crucial for avoiding the generation of harmful content. Existing research mechanistically interprets LLM safety behaviors at various perspectives to enhance transparency and improve safety. However, the association between attention heads—a core component for LLMs—and safety remains insufficiently understood, making it difficult to characterize the safety mechanisms of LLMs completely at a fine-grained level. In this work, we propose an attention head safety correlation detection and causal influence verification method, namely AHSCC, which identifies attention heads that control safety mechanisms in LLMs via representation similarity-based correlation detection and patch-based intervention causal influence verification. We find that numerous safety-related attention heads are distributed across the intermediate layers of LLMs, and by directly intervening on these attention heads, the safety behavior of LLMs can be significantly controlled. Therefore, we claim that LLM safety mechanisms are distributed to varying extents across numerous attention heads.Furthermore, building on this finding, we design an efficient attention head safety enhancement strategy named AHSFT, which further improves the safety of LLMs by lightweight fine-tuning specific safety-related attention heads. In summary, this work enhances the transparency of LLM safety, providing more comprehensive insights for interpreting and improving LLM safety mechanisms.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 25eab1a6-9707-4e28-820f-348a16151e4fRelated papers
- On the Role of Attention Heads in Large Language Model SafetyZhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu et al.ICLR 2025
- One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention HeadJunhao Xia, Haotian Zhu, Shuchao Pang, Zhigang Lu et al.NeurIPS 2025 · 5 citations
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai et al.NeurIPS 2025 · 53 citations
- Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language ModelsYanchen Yin, Dongqi Han, Linghui LiICML 2026 · 1 citation
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang et al.ICML 2026
