Lune

ICML2026顶会

Towards Understanding Massive Activations in Attention Sink Mechanism

Haiyu Wang, Yuanyuan Lin

出版方
2026年份

摘要

Recent studies have revealed two intriguing phenomena in large language models: attention sinks and massive activations. However, the co-emergence and co-existence of these two phenomena remain poorly understood.

In this work, we revisit the prevailing view that massive activations are the primary mechanism responsible for concentrating attention on sink tokens, and provide a more nuanced interpretation of their relationship. Through both theoretical analysis and empirical evidence, we demonstrate that massive activations and attention sinks jointly act to prevent excessive token mixing in selfattention. Specifically, attention sinks suppress mixing among non-sink tokens, whereas massive activations suppress mixing between sink tokens and non-sink tokens. Furthermore, our theory provides a principled explanation of how KVbiases, gating mechanisms and normalization layers can remove massive activations while largely preserving attention sinks. We further conduct intervention analyses and find that removing the value vector of the sink token can recover attention sinks even when massive activations are entirely suppressed. Overall, this work provides a mechanistic perspective on how massive activations and attention sinks interact under normalization and self-attention layers, offering new insights into their functional roles in Transformer models. Code is available at Attn- MA.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper31

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖