LayerMerge: Neural Network Depth Compression through Layer Pruning and Merging
Jinuk Kim, Marwa El Halabi, Mingi Ji, Hyun Oh Song
摘要
Recent works show that reducing the number of layers in a convolutional neural network can enhance efficiency while maintaining the performance of the network. Existing depth compression methods remove redundant non-linear activation functions and merge the consecutive convolution layers into a single layer. However, these methods suffer from a critical drawback; the kernel size of the merged layers becomes larger, significantly undermining the latency reduction gained from reducing the depth of the network. We show that this problem can be addressed by jointly pruning convolution layers and activation functions. To this end, we propose LayerMerge, a novel depth compression method that selects which activation layers and convolution layers to remove, to achieve a desired inference speedup while minimizing performance loss. Since the corresponding selection problem involves an exponential search space, we formulate a novel surrogate optimization problem and efficiently solve it via dynamic programming. Empirical results demonstrate that our method consistently outperforms existing depth compression and layer pruning methods on various network architectures, both on image classification and generation tasks. We release the code at https:// github.com/snu-mllab/LayerMerge .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion ModelsYoung D. Kwon, Rui Li, Sijia Li, Da Li 等AAAI 2026 · 被引用 5 次
- FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich TrainingFuhan Cai, Yong Guo, Jie Li, Wenbo Li 等AAAI 2026 · 被引用 2 次
- Transformer Block Coupling and its Correlation with Generalization in LLMsMurdock Aubry, Haoming Meng, Anton Sugolov, Vardan PapyanICLR 2025
- NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile DevicesRuchika Chavhan, Malcolm Chadwick, Alberto Gil Couto Pimentel Ramos, Luca Morreale 等ICML 2026
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
相关 Paper
- Efficient Latency-Aware CNN Depth Compression via Two-Stage Dynamic ProgrammingJinuk Kim, Yeonwoo Jeong, Deokjae Lee, Hyun Oh SongICML 2023 · 被引用 1 次
- Neuron Merging: Compensating for Pruned NeuronsWoojeong Kim, Suhyun Kim, Mincheol Park, Geunseok JeonNeurIPS 2020 · 被引用 42 次
- Convolutional Neural Network Pruning With Structural Redundancy ReductionZi Wang, Chengcheng Li, Xiangyang WangCVPR 2021
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-Based ApproachHaichuan Yang, Shupeng Gui, Yuhao Zhu, Ji LiuCVPR 2020
- GPTailor: Large Language Model Pruning Through Layer Cutting and StitchingGuinan Su, Li Shen, Lu Yin, Shiwei Liu 等ICLR 2026 · 被引用 3 次
