Advancing SVD-based LLM Compression via Layer-Wise Error Model Search
Moritz Thoma, Maximilian Groezinger, Maximilian Forstenhäusler, Emad Aghajanzadeh, Manoj Rohit Vemparala, Christos Anagnostopoulos, Pierpaolo Mori, Nael Fasfous, Alexander Frickenstein, Daniel Mueller-Gritschneder, Ulf Schlichtmann
Abstract
Low-rank SVD-based compression offers a powerful strategy to reduce the computational costs of LLMs. However, existing methods face two key limitations: (i) global rank allocation, where uncalibrated error proxies fail to capture complex error propagation, and (ii) decomposition quality, where Fisher-based estimators suffer from severe rank collapse. In this work, we address these limitations by introducing Layer-wise Error Modeling Search (LEMS) and KFAC-SVD. LEMS advances rank allocation by introducing a layer-wise error surrogate that integrates local and global layer importance alongside a propagation bias, enabling effective global rank allocation via an ILP formulation. KFAC-SVD improves decomposition quality by utilizing token-wise statistics, mitigating the rank deficiency observed in prior Fisher-based SVD approaches. Across Mistral, Qwen3, and Llama3 model families, we show that LEMS consistently outperforms existing search strategies, delivering significant zero-shot accuracy gains of up to 4.8 p.p. that generalize to model sizes of 70B parameters, while KFAC-SVD achieves an average perplexity improvement of 15%. Project Page & Code: https://lems-svd.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou et al.ICLR 2022 · 210 citations
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider et al.NeurIPS 2023 · 62 citations
- COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsJinqi Xiao, Miao Yin, Yu Gong, Xiao Zang et al.ICML 2023 · 17 citations
- Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model CompressionJingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li et al.ICLR 2025
Related papers
- SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM CompressionXing Hu, Dawei Yang, Yuan Cheng, Zhixuan Chen et al.ICLR 2026 · 9 citations
- CGSVD: Cascaded Granular Singular Value Decomposition for Large Language Model CompressionYuli Chen, Shuhao Zhang, Jiale Han, Fanshen Meng et al.ICML 2026
- Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM CompressionRuoling Qi, Yirui Liu, Xuaner Wu, Xiangyu Wang et al.ICML 2026
- A3: an Analytical Low-Rank Approximation Framework for AttentionJeffrey T. H. Wong, Cheng Zhang, Xinye Cao, Pedro Gimenes et al.ICML 2026 · 4 citations
- Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM CompressionAli Abbasi, Chayne Thrash, Haoran Qin, Shansita Sharma et al.ICML 2026 · 4 citations
