Towards Understanding Massive Activations in Attention Sink Mechanism
Haiyu Wang, Yuanyuan Lin
摘要
Recent studies have revealed two intriguing phenomena in large language models: attention sinks and massive activations. However, the co-emergence and co-existence of these two phenomena remain poorly understood.
In this work, we revisit the prevailing view that massive activations are the primary mechanism responsible for concentrating attention on sink tokens, and provide a more nuanced interpretation of their relationship. Through both theoretical analysis and empirical evidence, we demonstrate that massive activations and attention sinks jointly act to prevent excessive token mixing in selfattention. Specifically, attention sinks suppress mixing among non-sink tokens, whereas massive activations suppress mixing between sink tokens and non-sink tokens. Furthermore, our theory provides a principled explanation of how KVbiases, gating mechanisms and normalization layers can remove massive activations while largely preserving attention sinks. We further conduct intervention analyses and find that removing the value vector of the sink token can recover attention sinks even when massive activations are entirely suppressed. Overall, this work provides a mechanistic perspective on how massive activations and attention sinks interact under normalization and self-attention layers, offering new insights into their functional roles in Transformer models. Code is available at Attn- MA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang 等ICLR 2024 · 被引用 432 次
相关 Paper
- Anatomy of Massive Activations and Attention SinksShangwen Sun, Alfredo Canziani, Yann LeCun, Jiachen ZhuICML 2026
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same CoinEnrique Queipo-de-Llano, Alvaro Arroyo, Federico Barbero, Xiaowen Dong 等ICLR 2026 · 被引用 56 次
- A Single Layer to Explain Them All: Understanding Massive Values in Large Language ModelsZeru Shi, Zhenting Wang, Fan Yang, Qifan Wang 等ICML 2026
- The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension DisparitySiquan Li, Kaiqi Jiang, Jiacheng Sun, Tianyang HuICML 2026 · 被引用 1 次
- When Attention Sink Emerges in Language Models: An Empirical ViewXiangming Gu, Tianyu Pang, Chao Du, Qian Liu 等ICLR 2025
