Neuron Merging: Compensating for Pruned Neurons
Woojeong Kim, Suhyun Kim, Mincheol Park, Geunseok Jeon
Abstract
Network pruning is widely used to lighten and accelerate neural network models. Structured network pruning discards the whole neuron or filter, leading to accuracy loss. In this work, we propose a novel concept of neuron merging applicable to both fully connected layers and convolution layers, which compensates for the information loss due to the pruned neurons/filters. Neuron merging starts with decomposing the original weights into two matrices/tensors. One of them becomes the new weights for the current layer, and the other is what we name a scaling matrix, guiding the combination of neurons. If the activation function is ReLU, the scaling matrix can be absorbed into the next layer under certain conditions, compensating for the removed neurons. We also propose a data-free and inexpensive method to decompose the weights by utilizing the cosine similarity between neurons. Compared to the pruned model with the same topology, our merged model better preserves the output feature map of the original model; thus, it maintains the accuracy after pruning without fine-tuning. We demonstrate the effectiveness of our approach over network pruning for various model architectures and datasets. As an example, for VGG-16 on CIFAR-10, we achieve an accuracy of 93.16% while reducing 64% of total parameters, without any fine-tuning. The code can be found here: https://github.com/friendshipkim/neuron-merging
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bbb4937-9e65-45d7-a17c-696a33c3875bCited by top-tier papers7
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun et al.NeurIPS 2022 · 247 citations
- NOLA: Compressing LoRA using Linear Combination of Random BasisSoroush Abbasi Koohpayegani, Navaneet K. L., Parsa Nooralinejad, Soheil Kolouri et al.ICLR 2024 · 33 citations
- RED : Looking for Redundancies for Data-FreeStructured Compression of Deep Neural NetworksEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2021 · 32 citations
- Accurate Retraining-free Pruning for Pretrained Encoder-based Language ModelsSeungcheol Park, Hojun Choi, U KangICLR 2024 · 14 citations
- Balanced Column-Wise Block Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyun Jae Oh, Minkyu Kim et al.AAAI 2023 · 14 citations
Builds on2
Related papers
- : Cycle-Consistent Multi-Model MergingDonato Crisostomi, Marco Fumero, Daniele Baieri, Florian Bernard et al.NeurIPS 2024 · 23 citations
- Towards Meta-Pruning via Optimal TransportAlexander Theus, Olin Geimer, Friedrich Wicke, Thomas Hofmann et al.ICLR 2024 · 8 citations
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou et al.AAAI 2020 · 219 citations
- Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight AggregationFabian Morelli, Stephan EcksteinICML 2026
- LayerMerge: Neural Network Depth Compression through Layer Pruning and MergingJinuk Kim, Marwa El Halabi, Mingi Ji, Hyun Oh SongICML 2024 · 5 citations
