A Single Layer to Explain Them All: Understanding Massive Values in Large Language Models
Zeru Shi, Zhenting Wang, Fan Yang, Qifan Wang, Ruixiang Tang
Abstract
We investigate the origins of massive activations in large language models (LLMs) and identify a specific layer named the Massive Emergence Layer (ME Layer), that is consistently observed across model families, where massive activations first emerge and subsequently propagate to deeper layers through residual connections. We show that, within the ME Layer both the RMSNorm and the FFN parameters jointly contribute to the emergence of massive activations. Once formed, the massive activation token representation remains largely invariant across layers, reducing the diversity of hidden representations passed to the attention module. Motivated by this limitation, we propose a simple and effective method to reduce the rigidity of the massive activation token. Our approach consistently improves LLM performance across multiple tasks, including instruction following and math reasoning, in both training free and fine tuning settings. Moreover, we show that our method mitigates attention sinks by selectively weakening their influence, elucidating their origin at the hidden state level and shedding new light on principled mitigation strategies. The model and code have been released at MELayer & WeMask.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a699df4-7f36-40ae-8106-9780e3ff9e43Builds on24
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Anatomy of Massive Activations and Attention SinksShangwen Sun, Alfredo Canziani, Yann LeCun, Jiachen ZhuICML 2026
- Towards Understanding Massive Activations in Attention Sink MechanismHaiyu Wang, Yuanyuan LinICML 2026
- The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension DisparitySiquan Li, Kaiqi Jiang, Jiacheng Sun, Tianyang HuICML 2026 · 1 citation
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same CoinEnrique Queipo-de-Llano, Alvaro Arroyo, Federico Barbero, Xiaowen Dong et al.ICLR 2026 · 56 citations
- When Attention Sink Emerges in Language Models: An Empirical ViewXiangming Gu, Tianyu Pang, Chao Du, Qian Liu et al.ICLR 2025
