Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
Houwen Peng, Hao Du, Hongyuan Yu, Qi Li, Jing Liao, Jianlong Fu
摘要
One-shot weight sharing methods have recently drawn great attention in neural architecture search due to high efficiency and competitive performance. However, weight sharing across models has an inherent deficiency, i.e., insufficient training of subnetworks in hypernetworks. To alleviate this problem, we present a simple yet effective architecture distillation method. The central idea is that subnetworks can learn collaboratively and teach each other throughout the training process, aiming to boost the convergence of individual models. We introduce the concept of prioritized path, which refers to the architecture candidates exhibiting superior performance during training. Distilling knowledge from the prioritized paths is able to boost the training of subnetworks. Since the prioritized paths are changed on the fly depending on their performance and complexity, the final obtained paths are the cream of the crop. We directly select the most promising one from the prioritized paths as the final architecture, without using other complex search methods, such as reinforcement learning or evolution algorithms. The experiments on ImageNet verify such path distillation method can improve the convergence ratio and performance of the hypernetwork, as well as boosting the training of subnetworks. The discovered architectures achieve superior performance compared to the recent MobileNetV3 and EfficientNet families under aligned settings. Moreover, the experiments on object detection and more challenging search space show the generality and robustness of the proposed method. Code and models are available at https://github.com/microsoft/cream.git 2 . * Equal contribution. Work done when Hao and Hongyuan were interns at MSRA. † Corresponding authors. 2 We also provide another implementation based upon Microsoft NNI AutoML open source toolkit at here. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 被引用 335 次
- AutoSNN: Towards Energy-Efficient Spiking Neural NetworksByunggook Na, Jisoo Mok, Seongsik Park, Dongjin Lee 等ICML 2022 · 被引用 89 次
- Searching the Search Space of Vision TransformerMinghao Chen, Kan Wu, Bolin Ni, Houwen Peng 等NeurIPS 2021 · 被引用 74 次
- Wisdom of Committees: An Overlooked Approach To Faster and More Accurate ModelsXiaofang Wang, Dan Kondratyuk, Eric Christiansen, Kris M. Kitani 等ICLR 2022 · 被引用 61 次
- BatchQuant: Quantized-for-all Architecture Search with Robust QuantizerHaoping Bai, Meng Cao, Ping Huang, Jiulong ShanNeurIPS 2021 · 被引用 43 次
它引用的顶会 Paper12
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 被引用 725 次
- Evaluating The Search Phase of Neural Architecture SearchKaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat 等ICLR 2020 · 被引用 370 次
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 被引用 362 次
相关 Paper
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang 等CVPR 2022 · 被引用 9 次
- HEP-NAS: Towards Efficient Few-shot Neural Architecture Search via Hierarchical Edge PartitioningJianfeng Li, Jiawen Zhang, Feng Wang, Lianbo MaAAAI 2025
- AlphaNet: Improved Training of Supernets with Alpha-DivergenceDilin Wang, Chengyue Gong, Meng Li, Qiang Liu 等ICML 2021 · 被引用 52 次
- One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space ShrinkingMinghao Chen, Jianlong Fu, Haibin LingCVPR 2021
- PA&DA: Jointly Sampling PAth and DAta for Consistent NASShun Lu, Yu Hu, Longxing Yang, Zihao Sun 等CVPR 2023
