Go Wide, Then Narrow: Efficient Training of Deep Thin Networks
Denny Zhou, Mao Ye, Chen Chen, Tianjian Meng, Mingxing Tan, Xiaodan Song, Quoc V. Le, Qiang Liu, Dale Schuurmans
Abstract
For deploying a deep learning model into production, it needs to be both accurate and compact to meet the latency and memory constraints. This usually results in a network that is deep (to ensure performance) and yet thin (to improve computational efficiency). In this paper, we propose an efficient method to train a deep thin network with a theoretic guarantee. Our method is motivated by model compression. It consists of three stages. First, we sufficiently widen the deep thin network and train it until convergence. Then, we use this well-trained deep wide network to warm up (or initialize) the original deep thin network. This is achieved by layerwise imitation, that is, forcing the thin network to mimic the intermediate outputs of the wide network from layer to layer. Finally, we further fine tune this already well-initialized deep thin network. The theoretical guarantee is established by using the neural mean field analysis. It demonstrates the advantage of our layerwise imitation approach over backpropagation. We also conduct large-scale empirical experiments to validate the proposed method. By training with our method, ResNet50 can outperform ResNet101, and BERT BASE can be comparable with BERT LARGE , when ResNet101 and BERT LARGE are trained under the standard training procedures as in the literature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang et al.NeurIPS 2022 · 62 citations
- Network Augmentation for Tiny Deep LearningHan Cai, Chuang Gan, Ji Lin, Song HanICLR 2022 · 34 citations
- DepthShrinker: A New Compression Paradigm Towards Boosting Real-Hardware Efficiency of Compact Neural NetworksYonggan Fu, Haichuan Yang, Jiayi Yuan, Meng Li et al.ICML 2022 · 26 citations
- StackRec: Efficient Training of Very Deep Sequential Recommender Models by Iterative StackingJiachun Wang, Fajie Yuan, Jian Chen, Qingyao Wu et al.SIGIR 2021 · 25 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Good Subnetworks Provably Exist: Pruning via Greedy Forward SelectionMao Ye, Chengyue Gong, Lizhen Nie, Denny Zhou et al.ICML 2020 · 123 citations
Related papers
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song et al.KDD 2021 · 49 citations
- ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge SystemsShuo Huai, Lei Zhang, Di Liu, Weichen Liu et al.DAC 2021 · 16 citations
- Compressing Models with Few Samples: Mimicking then ReplacingHuanyu Wang, Junjie Liu, Xin Ma, Yang Yong et al.CVPR 2022 · 11 citations
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev et al.ICLR 2020 · 229 citations
