Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
Max Staats, Matthias Thamm, Bernd Rosenow
摘要
This work analyzes singular-value spectra of weight matrices in pretrained transformer models to understand how information is stored at both ends of the spectrum. Using Random Matrix Theory (RMT) as a zero information hypothesis, we associate agreement with RMT as evidence of randomness and deviations as evidence for learning. Surprisingly, we observe pronounced departures from RMT not only among the largest singular values -- the usual outliers -- but also among the smallest ones. A comparison of the associated singular vectors with the eigenvectors of the activation covariance matrices shows that there is considerable overlap wherever RMT is violated. Thus, significant directions in the data are captured by small singular values and their vectors as well as by the large ones. We confirm this empirically: zeroing out the singular values that deviate from RMT raises language-model perplexity far more than removing values from the bulk, and after fine-tuning the smallest decile can be the third most influential part of the spectrum. To explain how vectors linked to small singular values can carry more information than those linked to larger values, we propose a linear random-matrix model. Our findings highlight the overlooked importance of the low end of the spectrum and provide theoretical and practical guidance for SVD-based pruning and compression of large language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary LearningJunxuan Wang, Xuyang Ge, Wentao Shu, Zhengfu He 等ICML 2026 · 被引用 5 次
- Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling LawsFabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu 等ICML 2026 · 被引用 5 次
- Eigenvectors of Experts are Training-free Non-collapsing RoutersGiang Do, Hung Le, Truyen TranICML 2026 · 被引用 1 次
- Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor AttacksBethan Evans, Jared TannerICML 2026 · 被引用 1 次
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 被引用 1 次
它引用的顶会 Paper11
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov 等ICML 2024 · 被引用 295 次
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou 等ICLR 2022 · 被引用 210 次
相关 Paper
- Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM CompressionAli Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma 等ICML 2026 · 被引用 4 次
- LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language ModelsGuangyan Li, Yongqiang Tang, Wensheng ZhangICML 2024 · 被引用 11 次
- Numerical Optimizations for Weighted Low-rank Estimation on Language ModelsTing Hua, Yen-Chang Hsu, Felicity Wang, Qian Lou 等EMNLP 2022 · 被引用 7 次
- LoRA vs Full Fine-tuning: An Illusion of EquivalenceReece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha SharmaNeurIPS 2025 · 被引用 152 次
- SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model CompressionXin Wang, Yu Zheng, Zhongwei Wan, Mi ZhangICLR 2025 · 被引用 1 次
