Entropy-Based Block Pruning for Efficient Large Language Models
Liangwei Yang, Yuhui Xu, Juntao Tan, Doyen Sahoo, silvio savarese, Caiming Xiong, Huan Wang, Shelby Heinecke
摘要
As large language models continue to scale, their growing computational and storage demands pose significant challenges for real-world deployment. In this work, we investigate redundancy within Transformer-based models and propose an entropy-based pruning strategy to enhance efficiency while maintaining performance. Empirical analysis reveals that the entropy of hidden representations decreases in the early blocks but progressively increases across most subsequent blocks. This trend suggests that entropy serves as a more effective measure of information richness within computation blocks. Unlike cosine similarity, which primarily captures geometric relationships, entropy directly quantifies uncertainty and information content, making it a more reliable criterion for pruning. Extensive experiments demonstrate that our entropy-based pruning approach surpasses cosine similarity-based methods in reducing model size while preserving accuracy, offering a promising direction for efficient model deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WET: Mitigating World-Conditioned Knowledge Conflicts via World Entropy TetheringZixuan Wang, Yifei He, Zihan Wang, Kun Wang 等ICML 2026
- E²LoRA: Efficient and Effective Low-Rank Adaptation with Entropy-Guided Adaptive SharingMinglei Li, Peng Ye, Jingqi Ye, Haonan He 等ICLR 2026
它引用的顶会 Paper12
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
相关 Paper
- Rethinking Layer Relevance in Large Language Models Beyond Cosine SimilarityCristian Hinostroza, Rodrigo Toro Icarte, Christ Devia, Andres Carvallo 等ICLR 2026 · 被引用 4 次
- SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer BlocksJiwon Song, Kyungseok Oh, Taesu Kim, Hyungjun Kim 等ICML 2024 · 被引用 86 次
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic LensXixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li 等NeurIPS 2025 · 被引用 44 次
- Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language ModelsMax Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju 等ICML 2026 · 被引用 1 次
- Entroformer: A Transformer-based Entropy Model for Learned Image CompressionYichen Qian, Xiuyu Sun, Ming Lin, Zhiyu Tan 等ICLR 2022 · 被引用 194 次
