From Attention to Activation: Unraveling the Enigmas of Large Language Models
Prannay Kaul, Chengcheng Ma, Ismail Elezi, Jiankang Deng
摘要
We study two strange phenomena in auto-regressive Transformers: (1) the dominance of the first token in attention heads; (2) the occurrence of large outlier activations in the hidden states. We find that popular large language models, such as Llama attend maximally to the first token in 98% of attention heads, a behaviour we attribute to the softmax function. To mitigate this issue, we propose a reformulation of softmax to softmax-1. Furthermore, we identify adaptive optimisers, e.g., Adam, as the primary contributor to the large outlier activations and introduce OrthoAdam, a novel optimiser that utilises orthogonal matrices to transform gradients, to address this issue. Finally, not only do our methods prevent these phenomena from occurring, but additionally, they enable Transformers to sustain their performance when quantised using basic algorithms, something that standard methods are unable to do. In summary, our methods reduce the attention proportion on the first token from 65% to 3.3%, the activation kurtosis in the hidden states from 1657 to 3.1, and perplexity penalty under 4-bit weight quantisation from 3565 to 0.3. Code is available at https://github.com/prannaykaul/OrthoAdam.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingFeijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao 等ICLR 2026 · 被引用 17 次
- SinkTrack: Attention Sink based Context Anchoring for Large Language ModelsXu Liu, Guikun Chen, Wenguan WangICLR 2026 · 被引用 4 次
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersJonghyun Hong, Sungyoon LeeEMNLP 2025
- A Single Layer to Explain Them All: Understanding Massive Values in Large Language ModelsZeru Shi, Zhenting Wang, Fan Yang, Qifan Wang 等ICML 2026
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-PathYiheng Tao, Kaiwen Cheng, Yao Lu, Chang Liu 等ICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
相关 Paper
- Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language ModelsJungwoo Park, Taewhoo Lee, Chanwoong Yoon, Hyeon Hwang 等ACL 2025 · 被引用 7 次
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 被引用 196 次
- Systematic Outliers in Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang 等ICLR 2025 · 被引用 1 次
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag 等NeurIPS 2024 · 被引用 27 次
- Self-Adjust SoftmaxChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi 等EMNLP 2025
