Lune

NeurIPS2025顶会

ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization

Dmitriy Shopkhoev, Ammar Ali, Magauiya Zhussip, Valentin Malykh, Stamatios Lefkimmiatis, Nikos Komodakis, Sergey Zagoruyko

2025年份
10被引次数

摘要

We introduce ReplaceMe, a generalized training-free depth pruning method that effectively replaces transformer blocks with a linear operation, while maintaining high performance for low compression ratios. In contrast to conventional pruning approaches that require additional training or fine-tuning, our approach requires only a small calibration dataset that is used to estimate a linear transformation, which approximates the pruned blocks. The estimated linear mapping can be seamlessly merged with the remaining transformer blocks, eliminating the need for any additional network parameters. Our experiments show that ReplaceMe consistently outperforms other training-free approaches and remains highly competitive with state-of-the-art pruning methods that involve extensive retraining/fine-tuning and architectural modifications. Applied to several large language models (LLMs), ReplaceMe achieves up to 25% pruning while retaining approximately 90% of the original model's performance on open benchmarks-without any training or healing steps, resulting in minimal computational overhead. We provide an opensource library implementing ReplaceMe alongside several state-of-the-art depth pruning techniques, available at https://github.com/mts-ai/ReplaceMe.

enhances hardware utilization efficiency but also potentially achieves greater reductions in resource consumption. Importantly, it operates independently of the hardware type used.

In this work, we focus on structural depth pruning, operating under the hypothesis that a contiguous set of transformer blocks can be effectively approximated by a single linear transformation. To validate this idea, we propose ReplaceMe, a novel training-free pruning method that replaces selected blocks with a linear transformation estimated from a small calibration dataset. It should be noted that most existing pruning methods require a post-pruning retraining phase, often referred to as a "healing process", to recover lost performance. This retraining stage can be time-consuming and computationally expensive. In contrast, ReplaceMe preserves the majority of the model performance without any retraining for reasonable compression ratio scenarios. ReplaceMe generalizes depth pruning methods by introducing a simple yet effective linear transformation that compensates for the error caused by block removal. This transformation is subsequently fused with one of the remaining model weights, enabling seamless integration without adding parameters. The contributions of this work can be summarized as follows:

  1. We propose ReplaceMe, a generalized method for depth pruning that can maintain model performance without requiring any healing process, for reasonable compression ratios; 2. We conduct a detailed study on the estimation of the linear transformation with both analytical and numerical methods and under different objectives;

  2. We provide detailed ablation studies for different calibration data, solvers, and LLM architectures;

  3. We validate the effectiveness and generality of ReplaceMe across diverse model families, including LLMs and vision transformer architectures like ViT [7].

This paper is organized as follows: Section 2 presents the core methodology behind our trainingfree depth pruning approach. It introduces the framework for identifying prunable layers in large language models (LLMs) and estimating the corresponding linear transformations that compensate for the removed components. This section also discusses the selection of appropriate loss functions, regularization strategies to ensure generalizability, and the potential extension to multiple linear transformations for more flexible pruning. Section 3 then provides comprehensive experimental results and ablation studies, demonstrating the effectiveness and robustness of our method, and analyzing the key factors that influence its performance.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext bc39c47a-779b-4731-bd85-e1b689031310

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖