Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement Learning
Sixing Yu, Arya Mazaheri, Ali Jannesari
Abstract
Model compression is an essential technique for deploying deep neural networks (DNNs) on power and memory-constrained resources. However, existing model-compression methods often rely on human expertise and focus on parameters' local importance, ignoring the rich topology information within DNNs. In this paper, we propose a novel multi-stage graph embedding technique based on graph neural networks (GNNs) to identify DNN topologies and use reinforcement learning (RL) to find a suitable compression policy. We performed resource-constrained (i.e., FLOPs) channel pruning and compared our approach with state-of-the-art model compression methods. We evaluated our method on various models from typical to mobile-friendly networks, such as ResNet family, VGG-16, MobileNet-v1/v2, and Shuf-fleNet. Results show that our method can achieve higher compression ratios with a minimal finetuning cost yet yields outstanding and competitive performance. The code is open-sourced at https://github.com/yusx-swapp/ GNN-RL-Model-Compression .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Structural Alignment for Network Pruning through Partial RegularizationShangqian Gao, Zeyu Zhang, Yanfu Zhang, Feihu Huang et al.ICCV 2023 · 26 citations
- SPATL: Salient Parameter Aggregation and Transfer Learning for Heterogeneous Federated LearningSixing Yu, Phuong Nguyen, Waqwoya Abebe, Wei Qian et al.SC 2022 · 21 citations
- Auto- Train-Once: Controller Network Guided Automatic Network Pruning from ScratchXidong Wu, Shangqian Gao, Zeyu Zhang, Zhenzhen Li et al.CVPR 2024 · 13 citations
- Lossy and Lossless (L2) Post-training Model Size CompressionYumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong et al.ICCV 2023 · 5 citations
- V-Pruner: A Fast and Globally-informed Token Pruning Framework for Vision TransformerGuangzhen Yao, Jiayun Zheng, Zezhou Wang, Wenxin Zhang et al.AAAI 2026 · 1 citation
Builds on20
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression RatesNing Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang et al.AAAI 2020 · 204 citations
- Neuron-level Structured Pruning using Polarization RegularizerTao Zhuang, Zhixuan Zhang, Yuheng Huang, Xiaoyi Zeng et al.NeurIPS 2020 · 168 citations
- Provable Filter Pruning for Efficient Neural NetworksLucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman et al.ICLR 2020 · 161 citations
Related papers
- Auto Graph Encoder-Decoder for Neural Network PruningSixing Yu, Arya Mazaheri, Ali JannesariICCV 2021 · 47 citations
- DECORE: Deep Compression with Reinforcement LearningManoj Alwani, Yang Wang, Vashisht MadhavanCVPR 2022 · 42 citations
- Storage Efficient and Dynamic Flexible Runtime Channel Pruning via Deep Reinforcement LearningJianda Chen, Shangyu Chen, Sinno Jialin PanNeurIPS 2020 · 31 citations
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
- NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationZhengang Li, Geng Yuan, Wei Niu, Pu Zhao et al.CVPR 2021
