ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
Dmitriy Shopkhoev, Ammar Ali, Magauiya Zhussip, Valentin Malykh, Stamatios Lefkimmiatis, Nikos Komodakis, Sergey Zagoruyko
摘要
We introduce ReplaceMe, a generalized training-free depth pruning method that effectively replaces transformer blocks with a linear operation, while maintaining high performance for low compression ratios. In contrast to conventional pruning approaches that require additional training or fine-tuning, our approach requires only a small calibration dataset that is used to estimate a linear transformation, which approximates the pruned blocks. The estimated linear mapping can be seamlessly merged with the remaining transformer blocks, eliminating the need for any additional network parameters. Our experiments show that ReplaceMe consistently outperforms other training-free approaches and remains highly competitive with state-of-the-art pruning methods that involve extensive retraining/fine-tuning and architectural modifications. Applied to several large language models (LLMs), ReplaceMe achieves up to 25% pruning while retaining approximately 90% of the original model's performance on open benchmarks-without any training or healing steps, resulting in minimal computational overhead. We provide an opensource library implementing ReplaceMe alongside several state-of-the-art depth pruning techniques, available at https://github.com/mts-ai/ReplaceMe.
enhances hardware utilization efficiency but also potentially achieves greater reductions in resource consumption. Importantly, it operates independently of the hardware type used.
In this work, we focus on structural depth pruning, operating under the hypothesis that a contiguous set of transformer blocks can be effectively approximated by a single linear transformation. To validate this idea, we propose ReplaceMe, a novel training-free pruning method that replaces selected blocks with a linear transformation estimated from a small calibration dataset. It should be noted that most existing pruning methods require a post-pruning retraining phase, often referred to as a "healing process", to recover lost performance. This retraining stage can be time-consuming and computationally expensive. In contrast, ReplaceMe preserves the majority of the model performance without any retraining for reasonable compression ratio scenarios. ReplaceMe generalizes depth pruning methods by introducing a simple yet effective linear transformation that compensates for the error caused by block removal. This transformation is subsequently fused with one of the remaining model weights, enabling seamless integration without adding parameters. The contributions of this work can be summarized as follows:
-
We propose ReplaceMe, a generalized method for depth pruning that can maintain model performance without requiring any healing process, for reasonable compression ratios; 2. We conduct a detailed study on the estimation of the linear transformation with both analytical and numerical methods and under different objectives;
-
We provide detailed ablation studies for different calibration data, solvers, and LLM architectures;
-
We validate the effectiveness and generality of ReplaceMe across diverse model families, including LLMs and vision transformer architectures like ViT [7].
This paper is organized as follows: Section 2 presents the core methodology behind our trainingfree depth pruning approach. It introduces the framework for identifying prunable layers in large language models (LLMs) and estimating the corresponding linear transformations that compensate for the removed components. This section also discusses the selection of appropriate loss functions, regularization strategies to ensure generalizability, and the potential extension to multiple linear transformations for more flexible pruning. Section 3 then provides comprehensive experimental results and ablation studies, demonstrating the effectiveness and robustness of our method, and analyzing the key factors that influence its performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang 等AAAI 2024 · 被引用 130 次
- Search for Efficient Large Language ModelsXuan Shen, Pu Zhao, Yifan Gong, Zhenglun Kong 等NeurIPS 2024 · 被引用 23 次
- Streamlining Redundant Layers to Compress Large Language ModelsXiaodong Chen, Yuxuan Hu, Jing Zhang, Yanling Wang 等ICLR 2025 · 被引用 3 次
- Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude CompensationXinrui Chen, Hongxing Zhang, Fanyi Zeng, Yongxian Wei 等AAAI 2026 · 被引用 3 次
- Olica: Efficient Structured Pruning of Large Language Models without RetrainingJiujun He, Huazhen LinICML 2025
