LaCo: Layer-wise Compensation for Pruned Large Language Models
Yingen Liu, Fan Wu, Xuyan Pan, Ruihui Li, Zhuo Tang, Kenli Li
摘要
Pruning is essential for the efficient deployment of Large Language Models (LLMs); however, it causes severe performance degradation due to the structural distortion induced by sparsity. Existing recovery strategies, such as LoRA, predominantly employ global finetuning, often overlooking the mechanistic root of this degradation: the layer-wise accumulation and amplification of local errors. To address this limitation, we propose LaCo (Layerwise Compensation), a framework that reorients the recovery paradigm from global adaptation to hierarchical representation alignment. By sequentially optimizing each layer to reconstruct the model's hidden states, LaCo effectively intercepts the error propagation chain at its source. Extensive experiments demonstrate that LaCo surpasses parameter-efficient baselines in both perplexity reduction and zeroshot reasoning. Notably, it reduces recoverytime memory usage to approximately 1/7 of the baseline and requires only 2,048 unlabeled samples to match a LoRA model trained on 50k examples-achieving a ∼ 25× improvement in data efficiency. 0 5 10 15 20 25 30 Layer Index 10 5 10 4 10 3 10 2 10 1 10 0 10 1 MSE Distance (Log Scale) Layer-wise Error Accumulatio Intercepting Error Chain Impact of Error Propagation at 70% Sparsity Pruned (No Compensation) LaCo (Ours) Mitigated Error Accumulation (a) Error Accumulation 0 5 10 15 20 25 30 Layer Index
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 被引用 208 次
相关 Paper
- LSA: Layer-wise Sparsity Allocation for Large Language Model Pruning Based on Minimal Linear Reconstruction ErrorZhiguo Yang, Changjian Deng, Qinke Chen, Zijing Zhou 等ICLR 2026
- Dynamic Low-Rank Sparse Adaptation for Large Language ModelsWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Yang Liu 等ICLR 2025
- Restoring Pruned Large Language Models via Lost Component CompensationZijian Feng, Hanzhang Zhou, Zixiao Zhu, Tianjiao Li 等NeurIPS 2025 · 被引用 3 次
- Reassessing Layer Pruning in LLMs: New Insights and MethodsYao Lu, Hao Cheng, Yujie Fang, Zeyu Wang 等ICLR 2026 · 被引用 24 次
- Determining Layer-wise Sparsity for Large Language Models Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao 等ICML 2025
