LaCo: Layer-wise Compensation for Pruned Large Language Models
Yingen Liu, Fan Wu, Xuyan Pan, Ruihui Li, Zhuo Tang, Kenli Li
Abstract
Pruning is essential for the efficient deployment of Large Language Models (LLMs); however, it causes severe performance degradation due to the structural distortion induced by sparsity. Existing recovery strategies, such as LoRA, predominantly employ global finetuning, often overlooking the mechanistic root of this degradation: the layer-wise accumulation and amplification of local errors. To address this limitation, we propose LaCo (Layerwise Compensation), a framework that reorients the recovery paradigm from global adaptation to hierarchical representation alignment. By sequentially optimizing each layer to reconstruct the model's hidden states, LaCo effectively intercepts the error propagation chain at its source. Extensive experiments demonstrate that LaCo surpasses parameter-efficient baselines in both perplexity reduction and zeroshot reasoning. Notably, it reduces recoverytime memory usage to approximately 1/7 of the baseline and requires only 2,048 unlabeled samples to match a LoRA model trained on 50k examples-achieving a ∼ 25× improvement in data efficiency. 0 5 10 15 20 25 30 Layer Index 10 5 10 4 10 3 10 2 10 1 10 0 10 1 MSE Distance (Log Scale) Layer-wise Error Accumulatio Intercepting Error Chain Impact of Error Propagation at 70% Sparsity Pruned (No Compensation) LaCo (Ours) Mitigated Error Accumulation (a) Error Accumulation 0 5 10 15 20 25 30 Layer Index
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
Related papers
- LSA: Layer-wise Sparsity Allocation for Large Language Model Pruning Based on Minimal Linear Reconstruction ErrorZhiguo Yang, Changjian Deng, Qinke Chen, Zijing Zhou et al.ICLR 2026
- Dynamic Low-Rank Sparse Adaptation for Large Language ModelsWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Yang Liu et al.ICLR 2025
- Restoring Pruned Large Language Models via Lost Component CompensationZijian Feng, Hanzhang Zhou, Zixiao Zhu, Tianjiao Li et al.NeurIPS 2025 · 3 citations
- Reassessing Layer Pruning in LLMs: New Insights and MethodsYao Lu, Hao Cheng, Yujie Fang, Zeyu Wang et al.ICLR 2026 · 24 citations
- Determining Layer-wise Sparsity for Large Language Models Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao et al.ICML 2025
