Numerical Optimizations for Weighted Low-rank Estimation on Language Models
Ting Hua, Yen-Chang Hsu, Felicity Wang, Qian Lou, Yilin Shen, Hongxia Jin
Abstract
Singular value decomposition (SVD) is one of the most popular compression methods that approximate a target matrix with smaller matrices. However, standard SVD treats the parameters within the matrix with equal importance, which is a simple but unrealistic assumption. The parameters of a trained neural network model may affect the task performance unevenly, which suggests non-equal importance among the parameters. Compared to SVD, the decomposition method aware of parameter importance is the more practical choice in real cases. Unlike standard SVD, weighted value decomposition is a non-convex optimization problem that lacks a closed-form solution. We systematically investigated multiple optimization strategies to tackle the problem and examined our method by compressing Transformer-based language models. Further, we designed a metric to predict when the SVD may introduce a significant performance drop, for which our method can be a rescue strategy. The extensive evaluations demonstrate that our method can perform better than current SOTA methods in compressing Transformer-based language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82ca4a89-5b66-43aa-b438-cb594f8deb9aCited by top-tier papers2
- Reweighted Solutions for Weighted Low Rank ApproximationDavid P. Woodruff, Taisuke YasudaICML 2024 · 3 citations
- Direction Sensitivity-Based Knowledge Distillation: Optimization-Aware Low-Rank Knowledge TransferYongkai Liao, Xinxing Chen, Zhongzheng Fu, Haoyuan Wang et al.AAAI 2026
Builds on4
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou et al.ICLR 2022 · 210 citations
- Group Fisher Pruning for Practical Network CompressionLiyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou et al.ICML 2021 · 204 citations
Related papers
- ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMsYan Yang, Yixia Li, Hongru Wang, Xuetao Wei et al.ACL 2025 · 4 citations
- SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model CompressionXin Wang, Yu Zheng, Zhongwei Wan, Mi ZhangICLR 2025 · 1 citation
- IMPACT: Importance-Aware Activation Space ReconstructionMd Mokarram Chowdhury, Daniel Agyei Asante, Ernie Chang, Yang LiACL 2026 · 1 citation
- MoE-SVD: Structured Mixture-of-Experts LLMs Compression via Singular Value DecompositionWei Li, Lujun Li, Hao Gu, You-Liang Huang et al.ICML 2025
- AdaSVD: Singular Value Decomposition with Adaptive Mechanisms for Large Multimodal ModelsZhiteng Li, Mingyuan Xia, Jingyuan Zhang, Zheng Hui et al.CVPR 2026
