Spectral Filters, Dark Signals, and Attention Sinks
Nicola Cancedda
摘要
Projecting intermediate representations onto the vocabulary is an increasingly popular interpretation tool for transformer-based LLMs, also known as the logit lens (Nostalgebraist). We propose a quantitative extension to this approach and define spectral filters on intermediate representations based on partitioning the singular vectors of the vocabulary embedding and unembedding matrices into bands. We find that the signals exchanged in the tail end of the spectrum, i.e. corresponding to the singular vectors with smallest singular values, are responsible for attention sinking (Xiao et al., 2023) , of which we provide an explanation. We find that the negative log-likelihood of pretrained models can be kept low despite suppressing sizeable parts of the embedding spectrum in a layer-dependent way, as long as attention sinking is preserved. Finally, we discover that the representation of tokens that draw attention from many tokens have large projections on the tail end of the spectrum, and likely act as additional attention sinks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
- Confidence Regulation Neurons in Language ModelsAlessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov 等NeurIPS 2024 · 被引用 68 次
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same CoinEnrique Queipo-de-Llano, Alvaro Arroyo, Federico Barbero, Xiaowen Dong 等ICLR 2026 · 被引用 56 次
- Mitigating Overthinking in Large Reasoning Models via Manifold SteeringYao Huang, Huanran Chen, Shouwei Ruan, Yichi Zhang 等NeurIPS 2025 · 被引用 45 次
- Spectral Adapter: Fine-Tuning in Spectral SpaceFangzhao Zhang, Mert PilanciNeurIPS 2024 · 被引用 34 次
它引用的顶会 Paper11
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
- The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank ReductionPratyusha Sharma, Jordan T. Ash, Dipendra MisraICLR 2024 · 被引用 135 次
相关 Paper
- Improving Neural Language Generation with Spectrum ControlLingxiao Wang, Jing Huang, Kevin Huang, Ziniu Hu 等ICLR 2020 · 被引用 94 次
- Contribution Weights: A Geometrical Analysis of Self-Attention TransformersJake Cunningham, Nicola Muca CironeICML 2026
- What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language ModelsYingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu ShenCVPR 2026 · 被引用 6 次
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsSonghao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui 等KDD 2026 · 被引用 1 次
- Anatomy of Massive Activations and Attention SinksShangwen Sun, Alfredo Canziani, Yann LeCun, Jiachen ZhuICML 2026
