Neural Epitome Search for Architecture-Agnostic Network Compression
Daquan Zhou, Xiaojie Jin, Qibin Hou, Kaixin Wang, Jianchao Yang, Jiashi Feng
Abstract
Traditional compression methods including network pruning, quantization, low rank factorization and knowledge distillation all assume that network architectures and parameters are one-to-one mapped. In this work, we propose a new perspective on network compression, i.e., network parameters can be disentangled from the architectures. From this viewpoint, we present the Neural Epitome Search (NES), a new neural network compression approach that learns to find compact yet expressive epitomes for weight parameters of a specified network architecture end-to-end. The complete network to compress can be generated from the learned epitome via a novel transformation method that adaptively transforms the epitomes to match weight shapes of the given architecture. Compared with existing compression methods, NES allows the weight tensors to be independent of the architecture design and hence can achieve a good trade-off between model compression rate and performance given a specific model size constraint. Experiments demonstrate that, on ImageNet, when taking MobileNetV2 as backbone, our approach improves the full-model baseline by 1.47% in top-1 accuracy with 25% MAdd reduction, and with the same compression ratio, improves AutoML for Model Compression (AMC) by 2.5% in top-1 accuracy. Moreover, taking EfficientNet-B0 as baseline, our NES yields an improvement of 1.2% but has 10% less MAdd. In particular, our method achieves a new state-of-the-art results of 77.5% under mobile settings (<350M MAdd). Code will be made publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- TAda! Temporally-Adaptive Convolutions for Video UnderstandingZiyuan Huang, Shiwei Zhang, Liang Pan, Zhiwu Qing et al.ICLR 2022 · 72 citations
- DKM: Differentiable k-Means Clustering Layer for Neural Network CompressionMinsik Cho, Keivan Alizadeh-Vahid, Saurabh Adya, Mohammad RastegariICLR 2022 · 39 citations
- AutoSpace: Neural Architecture Search with Less Human InterferenceDaquan Zhou, Xiaojie Jin, Xiaochen Lian, Linjie Yang et al.ICCV 2021 · 11 citations
- EPIM: Efficient Processing-In-Memory Accelerators based on EpitomeChenyu Wang, Zhen Dong, Daquan Zhou, Zhenhua Zhu et al.DAC 2024
Related papers
- AutoGAN-Distiller: Searching to Compress Generative Adversarial NetworksYonggan Fu, Wuyang Chen, Haotao Wang, Haoran Li et al.ICML 2020 · 91 citations
- AutoMC: Automated Model Compression Based on Domain Knowledge and Progressive SearchChunnan Wang, Hongzhi Wang, Xiangyu ShiICDE 2024 · 2 citations
- AutoBSS: An Efficient Algorithm for Block Stacking Style SearchYikang Zhang, Jian Zhang, Zhao ZhongNeurIPS 2020 · 4 citations
- Auto Graph Encoder-Decoder for Neural Network PruningSixing Yu, Arya Mazaheri, Ali JannesariICCV 2021 · 47 citations
- Neural Architecture Search for Lightweight Non-Local NetworksYingwei Li, Xiaojie Jin, Jieru Mei, Xiaochen Lian et al.CVPR 2020
