Compressing Large Language Models using Low Rank and Low Precision Decomposition
Rajarshi Saha, Naomi Sagan, Varun Srivastava, Andrea Goldsmith, Mert Pilanci
Abstract
The prohibitive sizes of Large Language Models (LLMs) today make it difficult to deploy them on memory-constrained edge devices. This work introduces -- a new post-training LLM compression algorithm that harnesses the inherent low-rank structure of a weight matrix by approximating it via a low-rank, low-precision decomposition as . Here, and are low rank factors, and the entries of , and are quantized. The model is compressed by substituting each layer with its decomposition, and the zero-shot performance of the compressed model is evaluated. Additionally, and are readily amenable to low-rank adaptation, consequently enhancing the zero-shot performance. obtains this decomposition by formulating it as an optimization problem , where is the calibration data, and are constrained to be representable using low-precision formats. Theoretical upper bounds on the approximation error of are established using a rank-constrained regression framework, and the tradeoff between compression ratio and model performance is studied by analyzing the impact of target rank and quantization bit budget. Results illustrate that compressing LlaMa- B//B and LlaMa- B models using outperforms existing post-training LLM compression techniques in the regime of less than bits per parameter. The implementation is available at: https://github.com/pilancilab/caldera.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4302d13-5696-406e-b378-b4dc168fb6aeCited by top-tier papers25
- Small Singular Values Matter: A Random Matrix Analysis of Transformer ModelsMax Staats, Matthias Thamm, Bernd RosenowNeurIPS 2025 · 21 citations
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuningSirui Chen, Yunzhe Qi, Mengting Ai, Yifan Sun et al.ICLR 2026 · 9 citations
- LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal ModelsPengcheng Zheng, Chaoning Zhang, Jiarong Mo, Guohui Li et al.ICLR 2026 · 9 citations
- SALS: Sparse Attention in Latent Space for KV Cache CompressionJunlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu et al.NeurIPS 2025 · 7 citations
Builds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- A3: an Analytical Low-Rank Approximation Framework for AttentionJeffrey T. H. Wong, Cheng Zhang, Xinye Cao, Pedro Gimenes et al.ICML 2026 · 4 citations
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block SkippingYu-Chen Lu, Sheng-Feng Yu, Hui-Hsien Weng, Pei-Shuo Wang et al.AAAI 2026 · 1 citation
- LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model FinetuningHan Guo, Philip Greengard, Eric P. Xing, Yoon KimICLR 2024 · 94 citations
- QERA: an Analytical Framework for Quantization Error ReconstructionCheng Zhang, Jeffrey T. H. Wong, Can Xiao, George Anthony Constantinides et al.ICLR 2025
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
