Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions
boyu shi, Chang Liu, Chuanbao Gao, Xu Yang, Xin Geng
Abstract
Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representation-based analyses struggle to explain this mechanism. We propose studying pruning through decision representation. Focusing on multiple-choice tasks, we introduce two metrics, Decision Margin and Option Frequency, and an Iterative Pruning method to analyze layer-wise decision dynamics. Our findings reveal a sharp decision transition that partitions the network into two stages: a Silent Phase, where the model cannot yet predict the correct answer, and a Decisive Phase, where the correct prediction emerges. We also find that pruning the Decisive Phase has minimal impact, whereas pruning the Silent Phase triggers immediate performance collapse, highlighting its extreme sensitivity to structural changes. Therefore, we conclude that pruning-induced collapse stems from disrupting the Silent Phase, which prevents the critical decision transition from occurring.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a732766-19ca-46b8-bfc4-988bc43f7c47Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse ApproximationYixiao Li, Yifan Yu, Qingru Zhang, Chen Liang et al.ICML 2023 · 125 citations
- Learngene: From Open-World to Your Learning TaskQiu-Feng Wang, Xin Geng, Shuxia Lin, Shiyu Xia et al.AAAI 2022 · 37 citations
- Transformer as Linear Expansion of LearngeneShiyu Xia, Miaosen Zhang, Xu Yang, Ruiming Chen et al.AAAI 2024 · 14 citations
Related papers
- Demystifying When Pruning Works via Representation HierarchiesShwai He, Guoheng Sun, Haichao Zhang, Yun Fu et al.ICML 2026 · 2 citations
- Dual-Assessment Driven Pruning: Iterative Optimizing Layer-wise Sparsity for Large Language ModelQinghui Sun, Weilun Wang, Yanni Zhu, Shenghuan He et al.KDD 2024 · 3 citations
- Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMsChang Gao, Kang Zhao, Runqi Wang, Jianfei Chen et al.ACM MM 2025
- Persistent Topological Features in Large Language ModelsYuri Gardinazzi, Karthik Viswanathan, Giada Panerai, Alessio Ansuini et al.ICML 2025
- SkipGPT: Each Token is One of a KindAnhao Zhao, Fanghua Ye, Yingqi Fan, Junlong Tong et al.ICML 2025
