Neuron Merging: Compensating for Pruned Neurons
Woojeong Kim, Suhyun Kim, Mincheol Park, Geunseok Jeon
摘要
Network pruning is widely used to lighten and accelerate neural network models. Structured network pruning discards the whole neuron or filter, leading to accuracy loss. In this work, we propose a novel concept of neuron merging applicable to both fully connected layers and convolution layers, which compensates for the information loss due to the pruned neurons/filters. Neuron merging starts with decomposing the original weights into two matrices/tensors. One of them becomes the new weights for the current layer, and the other is what we name a scaling matrix, guiding the combination of neurons. If the activation function is ReLU, the scaling matrix can be absorbed into the next layer under certain conditions, compensating for the removed neurons. We also propose a data-free and inexpensive method to decompose the weights by utilizing the cosine similarity between neurons. Compared to the pruned model with the same topology, our merged model better preserves the output feature map of the original model; thus, it maintains the accuracy after pruning without fine-tuning. We demonstrate the effectiveness of our approach over network pruning for various model architectures and datasets. As an example, for VGG-16 on CIFAR-10, we achieve an accuracy of 93.16% while reducing 64% of total parameters, without any fine-tuning. The code can be found here: https://github.com/friendshipkim/neuron-merging
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun 等NeurIPS 2022 · 被引用 247 次
- NOLA: Compressing LoRA using Linear Combination of Random BasisSoroush Abbasi Koohpayegani, Navaneet K. L., Parsa Nooralinejad, Soheil Kolouri 等ICLR 2024 · 被引用 33 次
- RED : Looking for Redundancies for Data-FreeStructured Compression of Deep Neural NetworksEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2021 · 被引用 32 次
- Accurate Retraining-free Pruning for Pretrained Encoder-based Language ModelsSeungcheol Park, Hojun Choi, U KangICLR 2024 · 被引用 14 次
- Balanced Column-Wise Block Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyun Jae Oh, Minkyu Kim 等AAAI 2023 · 被引用 14 次
它引用的顶会 Paper2
相关 Paper
- : Cycle-Consistent Multi-Model MergingDonato Crisostomi, Marco Fumero, Daniele Baieri, Florian Bernard 等NeurIPS 2024 · 被引用 23 次
- Towards Meta-Pruning via Optimal TransportAlexander Theus, Olin Geimer, Friedrich Wicke, Thomas Hofmann 等ICLR 2024 · 被引用 8 次
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou 等AAAI 2020 · 被引用 219 次
- Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight AggregationFabian Morelli, Stephan EcksteinICML 2026
- LayerMerge: Neural Network Depth Compression through Layer Pruning and MergingJinuk Kim, Marwa El Halabi, Mingi Ji, Hyun Oh SongICML 2024 · 被引用 5 次
