Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
Max Staats, Matthias Thamm, Bernd Rosenow
Abstract
This work analyzes singular-value spectra of weight matrices in pretrained transformer models to understand how information is stored at both ends of the spectrum. Using Random Matrix Theory (RMT) as a zero information hypothesis, we associate agreement with RMT as evidence of randomness and deviations as evidence for learning. Surprisingly, we observe pronounced departures from RMT not only among the largest singular values -- the usual outliers -- but also among the smallest ones. A comparison of the associated singular vectors with the eigenvectors of the activation covariance matrices shows that there is considerable overlap wherever RMT is violated. Thus, significant directions in the data are captured by small singular values and their vectors as well as by the large ones. We confirm this empirically: zeroing out the singular values that deviate from RMT raises language-model perplexity far more than removing values from the bulk, and after fine-tuning the smallest decile can be the third most influential part of the spectrum. To explain how vectors linked to small singular values can carry more information than those linked to larger values, we propose a linear random-matrix model. Our findings highlight the overlooked importance of the low end of the spectrum and provide theoretical and practical guidance for SVD-based pruning and compression of large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dafe5036-d67d-4983-9878-d63fe7fe6a73Cited by top-tier papers9
- Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary LearningJunxuan Wang, Xuyang Ge, Wentao Shu, Zhengfu He et al.ICML 2026 · 5 citations
- Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling LawsFabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu et al.ICML 2026 · 5 citations
- Eigenvectors of Experts are Training-free Non-collapsing RoutersGiang Do, Hung Le, Truyen TranICML 2026 · 1 citation
- Theory of Minimal Weight Perturbations in Deep Networks and its Applications for Low-Rank Activated Backdoor AttacksBethan Evans, Jared TannerICML 2026 · 1 citation
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 1 citation
Builds on11
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov et al.ICML 2024 · 295 citations
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou et al.ICLR 2022 · 210 citations
Related papers
- Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM CompressionAli Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma et al.ICML 2026 · 4 citations
- LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language ModelsGuangyan Li, Yongqiang Tang, Wensheng ZhangICML 2024 · 11 citations
- Numerical Optimizations for Weighted Low-rank Estimation on Language ModelsTing Hua, Yen-Chang Hsu, Felicity Wang, Qian Lou et al.EMNLP 2022 · 7 citations
- LoRA vs Full Fine-tuning: An Illusion of EquivalenceReece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha SharmaNeurIPS 2025 · 152 citations
- SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model CompressionXin Wang, Yu Zheng, Zhongwei Wan, Mi ZhangICLR 2025 · 1 citation
