Auto Graph Encoder-Decoder for Neural Network Pruning
Sixing Yu, Arya Mazaheri, Ali Jannesari
Abstract
Model compression aims to deploy deep neural networks (DNN) on mobile devices with limited computing and storage resources. However, most of the existing model compression methods rely on manually defined rules, which require domain expertise. DNNs are essentially computational graphs, which contain rich structural information. In this paper, we aim to find a suitable compression policy from DNNs’ structural information. We propose an automatic graph encoder-decoder model compression (AGMC) method combined with graph neural networks (GNN) and reinforcement learning (RL). We model the target DNN as a graph and use GNN to learn the DNN’s embeddings automatically. We compared our method with rule-based DNN embedding model compression methods to show the effectiveness of our method. Results show that our learning-based DNN embedding achieves better performance and a higher compression ratio with fewer search steps. We evaluated our method on over-parameterized and mobile-friendly DNNs and compared our method with handcrafted and learning-based model compression approaches. On over parameterized DNNs, such as ResNet-56, our method outperformed handcrafted and learning-based methods with 4.36% and 2.56% higher accuracy, respectively. Furthermore, on MobileNet-v2, we achieved a higher compression ratio than state-of-the-art methods with just 0.93% accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Revisiting Random Channel Pruning for Neural Network CompressionYawei Li, Kamil Adamczewski, Wen Li, Shuhang Gu et al.CVPR 2022 · 114 citations
- Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement LearningSixing Yu, Arya Mazaheri, Ali JannesariICML 2022 · 54 citations
- SPATL: Salient Parameter Aggregation and Transfer Learning for Heterogeneous Federated LearningSixing Yu, Phuong Nguyen, Waqwoya Abebe, Wei Qian et al.SC 2022 · 21 citations
- Resource Constrained Model Compression via Minimax Optimization for Spiking Neural NetworksJue Chen, Huan Yuan, Jianchao Tan, Bin Chen et al.ACM MM 2023 · 5 citations
- Lossy and Lossless (L2) Post-training Model Size CompressionYumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong et al.ICCV 2023 · 5 citations
Builds on3
- BRP-NAS: Prediction-based NAS using GCNsLukasz Dudziak, Thomas Chau, Mohamed S. Abdelfattah, Royson Lee et al.NeurIPS 2020 · 233 citations
- AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression RatesNing Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang et al.AAAI 2020 · 204 citations
- Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONASHan Shi, Renjie Pi, Hang Xu, Zhenguo Li et al.NeurIPS 2020 · 148 citations
Related papers
- DECORE: Deep Compression with Reinforcement LearningManoj Alwani, Yang Wang, Vashisht MadhavanCVPR 2022 · 42 citations
- Neural Epitome Search for Architecture-Agnostic Network CompressionDaquan Zhou, Xiaojie Jin, Qibin Hou, Kaixin Wang et al.ICLR 2020 · 13 citations
- NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationZhengang Li, Geng Yuan, Wei Niu, Pu Zhao et al.CVPR 2021
- AutoMC: Automated Model Compression Based on Domain Knowledge and Progressive SearchChunnan Wang, Hongzhi Wang, Xiangyu ShiICDE 2024 · 2 citations
- Jointly Training and Pruning CNNs via Learnable Agent Guidance and AlignmentAlireza Ganjdanesh, Shangqian Gao, Heng HuangCVPR 2024
