Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
Houwen Peng, Hao Du, Hongyuan Yu, Qi Li, Jing Liao, Jianlong Fu
Abstract
One-shot weight sharing methods have recently drawn great attention in neural architecture search due to high efficiency and competitive performance. However, weight sharing across models has an inherent deficiency, i.e., insufficient training of subnetworks in hypernetworks. To alleviate this problem, we present a simple yet effective architecture distillation method. The central idea is that subnetworks can learn collaboratively and teach each other throughout the training process, aiming to boost the convergence of individual models. We introduce the concept of prioritized path, which refers to the architecture candidates exhibiting superior performance during training. Distilling knowledge from the prioritized paths is able to boost the training of subnetworks. Since the prioritized paths are changed on the fly depending on their performance and complexity, the final obtained paths are the cream of the crop. We directly select the most promising one from the prioritized paths as the final architecture, without using other complex search methods, such as reinforcement learning or evolution algorithms. The experiments on ImageNet verify such path distillation method can improve the convergence ratio and performance of the hypernetwork, as well as boosting the training of subnetworks. The discovered architectures achieve superior performance compared to the recent MobileNetV3 and EfficientNet families under aligned settings. Moreover, the experiments on object detection and more challenging search space show the generality and robustness of the proposed method. Code and models are available at https://github.com/microsoft/cream.git 2 . * Equal contribution. Work done when Hao and Hongyuan were interns at MSRA. † Corresponding authors. 2 We also provide another implementation based upon Microsoft NNI AutoML open source toolkit at here. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77d8f162-c353-43a0-aa26-be4d4e04ea1bCited by top-tier papers21
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 335 citations
- AutoSNN: Towards Energy-Efficient Spiking Neural NetworksByunggook Na, Jisoo Mok, Seongsik Park, Dongjin Lee et al.ICML 2022 · 89 citations
- Searching the Search Space of Vision TransformerMinghao Chen, Kan Wu, Bolin Ni, Houwen Peng et al.NeurIPS 2021 · 74 citations
- Wisdom of Committees: An Overlooked Approach To Faster and More Accurate ModelsXiaofang Wang, Dan Kondratyuk, Eric Christiansen, Kris M. Kitani et al.ICLR 2022 · 61 citations
- BatchQuant: Quantized-for-all Architecture Search with Robust QuantizerHaoping Bai, Meng Cao, Ping Huang, Jiulong ShanNeurIPS 2021 · 43 citations
Builds on12
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 725 citations
- Evaluating The Search Phase of Neural Architecture SearchKaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat et al.ICLR 2020 · 370 citations
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 362 citations
Related papers
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang et al.CVPR 2022 · 9 citations
- HEP-NAS: Towards Efficient Few-shot Neural Architecture Search via Hierarchical Edge PartitioningJianfeng Li, Jiawen Zhang, Feng Wang, Lianbo MaAAAI 2025
- AlphaNet: Improved Training of Supernets with Alpha-DivergenceDilin Wang, Chengyue Gong, Meng Li, Qiang Liu et al.ICML 2021 · 52 citations
- One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space ShrinkingMinghao Chen, Jianlong Fu, Haibin LingCVPR 2021
- PA&DA: Jointly Sampling PAth and DAta for Consistent NASShun Lu, Yu Hu, Longxing Yang, Zihao Sun et al.CVPR 2023
