Attention Heads Hold the Key to Understanding Safety Mechanisms in Large Language Models
Shiji Yang, Congyao Mei, Haodong Zou, Jie Chen, Shu Zhao
摘要
LLM safety mechanisms are crucial for avoiding the generation of harmful content. Existing research mechanistically interprets LLM safety behaviors at various perspectives to enhance transparency and improve safety. However, the association between attention heads—a core component for LLMs—and safety remains insufficiently understood, making it difficult to characterize the safety mechanisms of LLMs completely at a fine-grained level. In this work, we propose an attention head safety correlation detection and causal influence verification method, namely AHSCC, which identifies attention heads that control safety mechanisms in LLMs via representation similarity-based correlation detection and patch-based intervention causal influence verification. We find that numerous safety-related attention heads are distributed across the intermediate layers of LLMs, and by directly intervening on these attention heads, the safety behavior of LLMs can be significantly controlled. Therefore, we claim that LLM safety mechanisms are distributed to varying extents across numerous attention heads.Furthermore, building on this finding, we design an efficient attention head safety enhancement strategy named AHSFT, which further improves the safety of LLMs by lightweight fine-tuning specific safety-related attention heads. In summary, this work enhances the transparency of LLM safety, providing more comprehensive insights for interpreting and improving LLM safety mechanisms.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- On the Role of Attention Heads in Large Language Model SafetyZhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu 等ICLR 2025
- One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention HeadJunhao Xia, Haotian Zhu, Shuchao Pang, Zhigang Lu 等NeurIPS 2025 · 被引用 5 次
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai 等NeurIPS 2025 · 被引用 53 次
- Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language ModelsYanchen Yin, Dongqi Han, Linghui LiICML 2026 · 被引用 1 次
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang 等ICML 2026
