From Attention to Activation: Unraveling the Enigmas of Large Language Models
Prannay Kaul, Chengcheng Ma, Ismail Elezi, Jiankang Deng
Abstract
We study two strange phenomena in auto-regressive Transformers: (1) the dominance of the first token in attention heads; (2) the occurrence of large outlier activations in the hidden states. We find that popular large language models, such as Llama attend maximally to the first token in 98% of attention heads, a behaviour we attribute to the softmax function. To mitigate this issue, we propose a reformulation of softmax to softmax-1. Furthermore, we identify adaptive optimisers, e.g., Adam, as the primary contributor to the large outlier activations and introduce OrthoAdam, a novel optimiser that utilises orthogonal matrices to transform gradients, to address this issue. Finally, not only do our methods prevent these phenomena from occurring, but additionally, they enable Transformers to sustain their performance when quantised using basic algorithms, something that standard methods are unable to do. In summary, our methods reduce the attention proportion on the first token from 65% to 3.3%, the activation kurtosis in the hidden states from 1657 to 3.1, and perplexity penalty under 4-bit weight quantisation from 3565 to 0.3. Code is available at https://github.com/prannaykaul/OrthoAdam.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eda94f8-2707-4960-a3b6-3bbcbf13aec0Cited by top-tier papers5
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingFeijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao et al.ICLR 2026 · 17 citations
- SinkTrack: Attention Sink based Context Anchoring for Large Language ModelsXu Liu, Guikun Chen, Wenguan WangICLR 2026 · 4 citations
- Variance Sensitivity Induces Attention Entropy Collapse and Instability in TransformersJonghyun Hong, Sungyoon LeeEMNLP 2025
- A Single Layer to Explain Them All: Understanding Massive Values in Large Language ModelsZeru Shi, Zhenting Wang, Fan Yang, Qifan Wang et al.ICML 2026
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-PathYiheng Tao, Kaiwen Cheng, Yao Lu, Chang Liu et al.ICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language ModelsJungwoo Park, Taewhoo Lee, Chanwoong Yoon, Hyeon Hwang et al.ACL 2025 · 7 citations
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 196 citations
- Systematic Outliers in Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang et al.ICLR 2025 · 1 citation
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag et al.NeurIPS 2024 · 27 citations
- Self-Adjust SoftmaxChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi et al.EMNLP 2025
