Self-attention Networks Localize When QK-eigenspectrum Concentrates
Han Bao, Ryuichiro Hataya, Ryo Karakida
摘要
The self-attention mechanism prevails in modern machine learning. It has an interesting functionality of adaptively selecting tokens from an input sequence by modulating the degree of attention localization, which many researchers speculate is the basis of the powerful model performance but complicates the underlying mechanism of the learning dynamics. In recent years, mainly two arguments have connected attention localization to the model performances. One is the rank collapse, where the embedded tokens by a self-attention block become very similar across different tokens, leading to a less expressive network. The other is the entropy collapse, where the attention probability approaches non-uniform and entails low entropy, making the learning dynamics more likely to be trapped in plateaus. These two failure modes may apparently contradict each other because the rank and entropy collapses are relevant to uniform and non-uniform attention, respectively. To this end, we characterize the notion of attention localization by the eigenspectrum of query-key parameter matrices and reveal that a small eigenspectrum variance leads attention to be localized. Interestingly, the small eigenspectrum variance prevents both rank and entropy collapse, leading to better model expressivity and trainability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention DynamicsHugo Koubbi, Louis Hernandez, Matthieu BoussardICML 2026 · 被引用 7 次
- NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward NetworksNandan Kumar Jha, Brandon ReagenICLR 2026 · 被引用 4 次
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 被引用 1 次
- Intervening Anchor Token: Decoding Strategy in Alleviating Hallucinations for MLLMsFeilong Tang, Zile Huang, Chengzhi Liu, Qiang Sun 等ICLR 2025
- Faster Query-Key Learning Sharpens Attention in Self-Attention ModelsRahul Vashisht, Harish RamaswamyICML 2026
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
相关 Paper
- On the Role of Attention Masks and LayerNorm in TransformersXinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka 等NeurIPS 2024 · 被引用 54 次
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersJonghyun Hong, Sungyoon LeeEMNLP 2025
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto 等NeurIPS 2022 · 被引用 161 次
- Critical attention scaling in long-context transformersShi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe RigolletICLR 2026 · 被引用 22 次
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisationAlessio Giorlandino, Sebastian GoldtICLR 2026 · 被引用 15 次
