Demystifying Singular Defects in Large Language Models
Haoqi Wang, Tong Zhang, Mathieu Salzmann
摘要
Large transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the linear approximations of layers. However, in large language models (LLMs), the underlying causes of highnorm tokens remain largely unexplored, and their different properties from those of ViTs require a new analysis framework. In this paper, we provide both theoretical insights and empirical validation across a range of recent models, leading to the following observations: i) The layer-wise singular direction predicts the abrupt explosion of token norms in LLMs. ii) The negative eigenvalues of a layer explain its sudden decay. iii) The computational pathways leading to high-norm tokens differ between initial and noninitial tokens. iv) High-norm tokens are triggered by the right leading singular vector of the matrix approximating the corresponding modules. We showcase two practical applications of these findings: the improvement of quantization schemes and the design of LLM signatures. Our findings not only advance the understanding of singular defects in LLMs but also open new avenues for their application. We expect that this work will stimulate further research into the internal mechanisms of LLMs. Code is released at https://github. com/haoqiwang/singular_defect .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Neural LineageRunpeng Yu, Xinchao WangCVPR 2024
相关 Paper
- A Single Layer to Explain Them All: Understanding Massive Values in Large Language ModelsZeru Shi, Zhenting Wang, Fan Yang, Qifan Wang 等ICML 2026
- Confidence Regulation Neurons in Language ModelsAlessandro Stolfo, Ben Wu, Wes Gurnee, Yonatan Belinkov 等NeurIPS 2024 · 被引用 68 次
- Transformers need glasses! Information over-squashing in language tasksFederico Barbero, Andrea Banino, Steven Kapturowski, Dharshan Kumaran 等NeurIPS 2024 · 被引用 105 次
- Transformer Block Coupling and its Correlation with Generalization in LLMsMurdock Aubry, Haoming Meng, Anton Sugolov, Vardan PapyanICLR 2025
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language ModelsJiayun Luo, Wan-Cyuan Fan, Lyuyang Wang, Xiangteng He 等ICLR 2026 · 被引用 19 次
