Compressing LLMs: The Truth is Rarely Pure and Never Simple
Ajay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang, Zhangyang Wang, Yinfei Yang
摘要
Despite their remarkable achievements, modern Large Language Models (LLMs) face exorbitant computational and memory footprints. Recently, several works have shown significant success in training-free and data-free compression (pruning and quantization) of LLMs that achieve 50 - 60% sparsity and reduce the bit width to 3 or 4 bits per weight, with negligible degradation of perplexity over the uncompressed baseline. As recent research efforts are focused on developing increasingly sophisticated compression methods, our work takes a step back and re-evaluates the effectiveness of existing SoTA compression methods, which rely on a fairly simple and widely questioned metric, perplexity (even for dense LLMs). We introduce Knowledge-Intensive Compressed LLM BenchmarK (LLM-KICK), a collection of carefully curated tasks to redefine the evaluation protocol for compressed LLMs, which have significant alignment with their dense counterparts and perplexity fail to capture subtle change in their true capabilities. LLM-KICK unveils many favorable merits and unfortunate plights of current SoTA compression methods: all pruning methods suffer significant performance degradation, sometimes at trivial sparsity ratios (e.g., 25-30%), and fail for N:M sparsity in knowledge-intensive tasks; current quantization methods are more successful than pruning; yet, pruned LLMs even at % sparsity are robust in-context retrieval and summarization systems; among others. LLM-KICK is designed to holistically access compressed LLMs' ability for language understanding, reasoning, generation, in-context retrieval, in-context summarization, etc. We hope our study can foster the development of better LLM compression methods. The reproduced codes are available at https://github.com/VITA-Group/llm-kick.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High SparsityLu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh 等ICML 2024 · 被引用 183 次
- LLM in a flash: Efficient Large Language Model Inference with Limited MemoryKeivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard 等ACL 2024 · 被引用 73 次
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionJunyuan Hong, Jinhao Duan, Chenhui Zhang, Zhangheng Li 等ICML 2024 · 被引用 54 次
- Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsLujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu 等NeurIPS 2024 · 被引用 51 次
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compressionMike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie 等ICLR 2026 · 被引用 47 次
它引用的顶会 Paper26
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
相关 Paper
- LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitChengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong 等AAAI 2026 · 被引用 3 次
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM CompressionPeijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li 等ICML 2025
- Accuracy is Not All You NeedAbhinav Dutta, Sanjeev Krishnan, Nipun Kwatra, Ramachandran RamjeeNeurIPS 2024 · 被引用 34 次
- Accurate Retraining-free Pruning for Pretrained Encoder-based Language ModelsSeungcheol Park, Hojun Choi, U KangICLR 2024 · 被引用 14 次
- XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer CompressionHaoqi Yang, Yao Yao, Zuchao Li, Baoyuan Qi 等EMNLP 2025
