SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
Wei Huang, Haotong Qin, Yangdong Liu, Yawei Li, Qinshuo Liu, Xianglong Liu, Luca Benini, Michele Magno, Shiming Zhang, Xiaojuan Qi
摘要
Post-training quantization (PTQ) is an effective technique for compressing large language models (LLMs). However, while uniform-precision quantization is computationally efficient, it often compromises model performance. To address this, we propose SliM-LLM, a saliencedriven mixed-precision quantization framework that allocates bit-widths at the group-wise. Our approach leverages the observation that important weights follow a structured distribution and introduces two key components: 1) Salience-Determined Bit Allocation adaptively assigns bitwidths to groups within each layer based on their salience; and 2) Salience-Weighted Quantizer Calibration optimizes quantizer parameters by incorporating element-level salience. With its structured partitioning, SliM-LLM provides a hardware-friendly solution that matches the efficiency of uniform quantization methods while improving accuracy. Experiments show that SliM-LLM achieves superior performance across various LLMs at low bit-widths. For example, a 2-bit quantized LLaMA-7B model reduces memory usage by nearly 6x compared to the floating-point baseline, decreases perplexity by 48% compared to state-of-the-art gradient-free PTQ methods, and maintains GPU inference speed. Additionally, the extended version, SliM-LLM + , which incorporates gradient-based quantization, further reduces perplexity by 35.1%. Our code is available at https://github.com/Aaronhuang-778/SliM-LLM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMsWei Huang, Yi Ge, Shuai Yang, Yicheng Xiao 等ICLR 2026 · 被引用 19 次
- FPTQuant: Function-Preserving Transforms for LLM QuantizationBoris van Breugel, Yelysei Bondarenko, Paul Whatmough, Markus NagelICML 2026 · 被引用 13 次
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language ModelsTianao Zhang, Zhiteng Li, Xianglong Yan, Haotong Qin 等ICLR 2026 · 被引用 11 次
- DecDEC: A Systems Approach to Advancing Low-Bit LLM QuantizationYeonhong Park, Jake Hyun, Hojoon Kim, Jae W. LeeOSDI 2025 · 被引用 9 次
- AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-TuningChanghai Zhou, Shiyang Zhang, Yuhua Zhou, Qian Qiao 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
相关 Paper
- MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationChun Hu, Junhui He, Shangyu Wu, Yuxin He 等EMNLP 2025 · 被引用 1 次
- SliderQuant: Accurate Post-Training Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan 等ICLR 2026 · 被引用 6 次
- BiLLM: Pushing the Limit of Post-Training Quantization for LLMsWei Huang, Yangdong Liu, Haotong Qin, Ying Li 等ICML 2024 · 被引用 161 次
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsChao Zeng, Songwei Liu, Yusheng Xie, Hong Liu 等AAAI 2025 · 被引用 24 次
- SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language ModelsHan Liu, Haotian Gao, Xiaotong Zhang, Changya Li 等KDD 2025 · 被引用 1 次
