SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
Mohammad Mozaffari, Amir Yazdanbakhsh, Maryam Mehri Dehnavi
Abstract
Conventional model compression techniques for LLMs address high memory consumption and slow inference challenges but typically require computationally expensive retraining to preserve accuracy. In contrast, one-shot compression methods eliminate retraining costs, but struggle to achieve accuracy comparable to dense models. This paper presents SLIM, a new one-shot compression framework that holistically integrates hardware-friendly quantization, sparsity, and lowrank approximation into a unified process. First, we formulate the quantization process using a probabilistic approach (SLIM-Quant) that enables us to apply uniform quantization. Then, we use an existing one-shot pruning method to apply semi-structured sparsity on top of the quantized weights. Finally, to compensate for the introduced aggregated quantization and sparsity error, we use a novel saliency function with unique invertible and additive features that enables us to mathematically compute the value of low-rank adapters. SLIM improves model accuracy by up to 5.66% (LLaMA-2-7B) for 2:4 sparsity with 4-bit weight quantization, outperforming prior methods. Models compressed with SLIM achieve up to 4.3× and 3.8× layer-wise speedup on Nvidia RTX3060 and A100 GPUs, respectively. Additionally, they achieve up to 0.23× end-to-end memory reduction in comparison to their dense counterparts. We also propose an optional PEFT recipe that further improves accuracy by up to 1.66% (LLaMA-2-13B) compared to SLIM without fine-tuning. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84eee3d1-8023-46c2-b296-561d687b6edcCited by top-tier papers4
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 12 citations
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin et al.ICLR 2026 · 6 citations
- Large Language Model Compression with Global Rank and Sparsity OptimizationChanghai Zhou, Qian Qiao, Yuhua Zhou, Yuxin Wu et al.ICLR 2026 · 6 citations
- RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion ModelsXing Cong, Hanlin Tang, Kan Liu, Lan Tao et al.ICML 2026 · 1 citation
Builds on15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
Related papers
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language ModelsWei Huang, Haotong Qin, Yangdong Liu, Yawei Li et al.ICML 2025
- OneBit: Towards Extremely Low-bit Large Language ModelsYuzhuang Xu, Xu Han, Zonghan Yang, Shuo Wang et al.NeurIPS 2024 · 110 citations
- Compress Large Language Models via Collaboration Between Learning and Matrix ApproximationYuesen Liao, Zhiwei Li, Binrui Wu, Zihao Cheng et al.NeurIPS 2025 · 1 citation
- WRP: Weight Recover Prune for Structured SparsityZhendong Tan, Xingjun Zhang, Zheng WeiACL 2024
- Pruning Large Language Models with Semi-Structural Adaptive Sparse TrainingWeiyu Huang, Yuezhou Hu, Guohao Jian, Jun Zhu et al.AAAI 2025 · 25 citations
