Lune

KDD2026顶会

Attention Heads Hold the Key to Understanding Safety Mechanisms in Large Language Models

Shiji Yang, Congyao Mei, Haodong Zou, Jie Chen, Shu Zhao

2026年份

摘要

LLM safety mechanisms are crucial for avoiding the generation of harmful content. Existing research mechanistically interprets LLM safety behaviors at various perspectives to enhance transparency and improve safety. However, the association between attention heads—a core component for LLMs—and safety remains insufficiently understood, making it difficult to characterize the safety mechanisms of LLMs completely at a fine-grained level. In this work, we propose an attention head safety correlation detection and causal influence verification method, namely AHSCC, which identifies attention heads that control safety mechanisms in LLMs via representation similarity-based correlation detection and patch-based intervention causal influence verification. We find that numerous safety-related attention heads are distributed across the intermediate layers of LLMs, and by directly intervening on these attention heads, the safety behavior of LLMs can be significantly controlled. Therefore, we claim that LLM safety mechanisms are distributed to varying extents across numerous attention heads.Furthermore, building on this finding, we design an efficient attention head safety enhancement strategy named AHSFT, which further improves the safety of LLMs by lightweight fine-tuning specific safety-related attention heads. In summary, this work enhances the transparency of LLM safety, providing more comprehensive insights for interpreting and improving LLM safety mechanisms.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖