Entropy-Based Block Pruning for Efficient Large Language Models
Liangwei Yang, Yuhui Xu, Juntao Tan, Doyen Sahoo, silvio savarese, Caiming Xiong, Huan Wang, Shelby Heinecke
Abstract
As large language models continue to scale, their growing computational and storage demands pose significant challenges for real-world deployment. In this work, we investigate redundancy within Transformer-based models and propose an entropy-based pruning strategy to enhance efficiency while maintaining performance. Empirical analysis reveals that the entropy of hidden representations decreases in the early blocks but progressively increases across most subsequent blocks. This trend suggests that entropy serves as a more effective measure of information richness within computation blocks. Unlike cosine similarity, which primarily captures geometric relationships, entropy directly quantifies uncertainty and information content, making it a more reliable criterion for pruning. Extensive experiments demonstrate that our entropy-based pruning approach surpasses cosine similarity-based methods in reducing model size while preserving accuracy, offering a promising direction for efficient model deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- WET: Mitigating World-Conditioned Knowledge Conflicts via World Entropy TetheringZixuan Wang, Yifei He, Zihan Wang, Kun Wang et al.ICML 2026
- E²LoRA: Efficient and Effective Low-Rank Adaptation with Entropy-Guided Adaptive SharingMinglei Li, Peng Ye, Jingqi Ye, Haonan He et al.ICLR 2026
Builds on12
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
Related papers
- Rethinking Layer Relevance in Large Language Models Beyond Cosine SimilarityCristian Hinostroza, Rodrigo Toro Icarte, Christ Devia, Andres Carvallo et al.ICLR 2026 · 4 citations
- SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer BlocksJiwon Song, Kyungseok Oh, Taesu Kim, Hyungjun Kim et al.ICML 2024 · 86 citations
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic LensXixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li et al.NeurIPS 2025 · 44 citations
- Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language ModelsMax Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju et al.ICML 2026 · 1 citation
- Entroformer: A Transformer-based Entropy Model for Learned Image CompressionYichen Qian, Xiuyu Sun, Ming Lin, Zhiyu Tan et al.ICLR 2022 · 194 citations
