Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
Qinsi Wang, Jinghan Ke, Masayoshi Tomizuka, Kurt Keutzer, Chenfeng Xu
Abstract
Large language models (LLMs) have sparked a new wave of AI applications; however, their substantial computational costs and memory demands pose significant challenges to democratizing access to LLMs for a broader audience. Singular Value Decomposition (SVD), a technique studied for decades, offers a hardwareindependent and flexibly tunable solution for LLM compression. In this paper, we present new directions using SVD: we first theoretically and experimentally analyze the optimality of directly truncating activations, then we further identify three key issues on SVD-based LLM compression, including (1) How can we determine the optimal truncation position for each layer in LLMs? (2) How can we efficiently update the weight matrices based on truncated activations? (3) How can we address the inherent "injection" nature that results in the information loss of the SVD? We propose a new paradigm for SVD-based LLM compression, Dobi-SVD, to tackle the three issues. First, we propose a differentiable truncation mechanism, along with gradient-robust backpropagation, enabling the model to adaptively find the optimal truncation positions. Next, we utilize the Eckart-Young-Mirsky theorem to derive a theoretically optimal weight update formula through rigorous mathematical analysis. Lastly, by observing and leveraging the quantization-friendly nature of matrices after SVD, we reconstruct a mapping between truncation positions and memory requirements, establishing a bijection from truncation positions to memory. Experimental results show that with a 40% parameter-compression rate, our method achieves a perplexity of 9.07 on the Wikitext2 dataset with the compressed LLama-7B model, a 78.7% improvement over the state-of-the-art SVD for LLM compression method. We emphasize that Dobi-SVD is the first to achieve such a high-ratio LLM compression while maintaining competitive performance. We also extend our Dobi-SVD to vision-language models (VLMs) and vision-languageaction models (VLAs), thereby highlighting its generalizability and practical value. We hope that the inference speedup-up to 12.4x on 12GB NVIDIA Titan Xp GPUs and 3x on 80GB A100 GPUs for LLMs, 1.2x and 1.17x on 80GB A100 GPUs for VLMs and VLAs, respectively -will bring significant benefits to the broader community such as multi-modal learning and robotics etc. * Qinsi Wang and Jinghan Ke contribute equally. Order is decided by coin flip. This work was done when Jinghan Ke visited UC Berkeley. † Chenfeng Xu advises the work and is the corresponding author.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb00db83-4040-42ff-b4d9-401e930cd262Cited by top-tier papers18
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent SystemsHancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang et al.NeurIPS 2025 · 42 citations
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language ModelsYutong Wang, Haiyu Wang, Sai Qian ZhangNeurIPS 2025 · 16 citations
- Angles Don't Lie: Unlocking Training‑Efficient RL Through the Model's Own SignalsQinsi Wang, Jinghan Ke, Hancheng Ye, Yueqian Lin et al.NeurIPS 2025 · 16 citations
- SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM CompressionXing Hu, Dawei Yang, Yuan Cheng, Zhixuan Chen et al.ICLR 2026 · 9 citations
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin et al.ICLR 2026 · 6 citations
Builds on18
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model CompressionXin Wang, Yu Zheng, Zhongwei Wan, Mi ZhangICLR 2025 · 1 citation
- Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM CompressionRuoling Qi, Yirui Liu, Xuaner Wu, Xiangyu Wang et al.ICML 2026
- Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model CompressionJingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li et al.ICLR 2025
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank ModelsZishan Shao, Yixiao Wang, Qinsi Wang, Ting Jiang et al.AAAI 2026 · 2 citations
- CGSVD: Cascaded Granular Singular Value Decomposition for Large Language Model CompressionYuli Chen, Shuhao Zhang, Jiale Han, Fanshen Meng et al.ICML 2026
