What are you sinking? A geometric approach on attention sink
Valeria Ruscio, Umberto Nanni, Fabrizio Silvestri
Abstract
Attention sink (AS) is a consistent pattern in transformer attention maps where certain tokens (often special tokens or positional anchors) disproportionately attract attention from other tokens. We show that in transformers, AS is not an architectural artifact, but it is the manifestation of a fundamental geometric principle: the establishment of reference frames that anchor representational spaces. We analyze several architectures and identify three distinct reference frame types, centralized, distributed, and bidirectional, that correlate with the attention sink phenomenon. We show that they emerge during the earliest stages of training as optimal solutions to the problem of establishing stable coordinate systems in high-dimensional spaces. We show the influence of architecture components, particularly position encoding implementations, on the specific type of reference frame. This perspective transforms our understanding of transformer attention mechanisms and provides insights for both architecture design and the relationship with AS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 649ff94d-74f2-4d2e-9941-609045b3b687Cited by top-tier papers7
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo et al.ICML 2026 · 5 citations
- SinkTrack: Attention Sink based Context Anchoring for Large Language ModelsXu Liu, Guikun Chen, Wenguan WangICLR 2026 · 4 citations
- NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward NetworksNandan Kumar Jha, Brandon ReagenICLR 2026 · 4 citations
- Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context FocusingLingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong et al.ACL 2026 · 3 citations
- Affine-Scaled Attention: Towards Flexible and Stable Transformer AttentionJeongin Bae, baeseong park, Gunho Park, Minsub Kim et al.ICML 2026 · 1 citation
Builds on13
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Topological AutoencodersMichael Moor, Max Horn, Bastian Rieck, Karsten M. BorgwardtICML 2020 · 192 citations
Related papers
- Anatomy of Massive Activations and Attention SinksShangwen Sun, Alfredo Canziani, Yann LeCun, Jiachen ZhuICML 2026
- When Attention Sink Emerges in Language Models: An Empirical ViewXiangming Gu, Tianyu Pang, Chao Du, Qian Liu et al.ICLR 2025
- The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension DisparitySiquan Li, Kaiqi Jiang, Jiacheng Sun, Tianyang HuICML 2026 · 1 citation
- Frayed RoPE and Long Inputs: A Geometric PerspectiveDavis Wertheimer, Aozhong Zhang, Derrick Liu, Penghang Yin et al.ICLR 2026 · 3 citations
- Absolute Position Embedding Learns Sinusoid-like Waves for Attention Based on Relative PositionYuji Yamamoto, Takuya MatsuzakiEMNLP 2023 · 1 citation
