AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling
Dilin Wang, Meng Li, Chengyue Gong, Vikas Chandra
Abstract
Neural architecture search (NAS) has shown great promise in designing state-of-the-art (SOTA) models that are both accurate and efficient. Recently, two-stage NAS, e.g. BigNAS, decouples the model training and searching process and achieves remarkable search efficiency and accuracy. Two-stage NAS requires sampling from the search space during training, which directly impacts the accuracy of the final searched models. While uniform sampling has been widely used for its simplicity, it is agnostic of the model performance Pareto front, which is the main focus in the search process, and thus, misses opportunities to further improve the model accuracy. In this work, we propose At-tentiveNAS that focuses on improving the sampling strategy to achieve better performance Pareto. We also propose algorithms to efficiently and effectively identify the networks on the Pareto during training. Without extra re-training or post-processing, we can simultaneously obtain a large number of networks across a wide range of FLOPs. Our discovered model family, AttentiveNAS models, achieves top-1 accuracy from 77.3% to 80.7% on ImageNet, and outperforms SOTA models, including BigNAS, Once-for-All networks and FBNetV3. We also achieve ImageNet accuracy of 80.1% with only 491 MFLOPs. Our training code and pretrained models are available at https://github . com/facebookresearch/AttentiveNAS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2e0ddd7-b4b5-4971-8e58-2c1a1c8d0102Cited by top-tier papers27
- Searching the Search Space of Vision TransformerMinghao Chen, Kan Wu, Bolin Ni, Houwen Peng et al.NeurIPS 2021 · 74 citations
- AlphaNet: Improved Training of Supernets with Alpha-DivergenceDilin Wang, Chengyue Gong, Meng Li, Qiang Liu et al.ICML 2021 · 52 citations
- ATPFL: Automatic Trajectory Prediction Model Design under Federated Learning FrameworkChunnan Wang, Xiang Chen, Junzhe Wang, Hongzhi WangCVPR 2022 · 41 citations
- CompOFA - Compound Once-For-All Networks for Faster Multi-Platform DeploymentManas Sahni, Shreya Varshini, Alind Khare, Alexey TumanovICLR 2021 · 37 citations
- Generalized Global Ranking-Aware Neural Architecture Ranker for Efficient Image Classifier SearchBicheng Guo, Tao Chen, Shibo He, Haoyu Liu et al.ACM MM 2022 · 21 citations
Builds on7
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 362 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
Related papers
- FP-NAS: Fast Probabilistic Neural Architecture SearchZhicheng Yan, Xiaoliang Dai, Peizhao Zhang, Yuandong Tian et al.CVPR 2021
- FBNetV3: Joint Architecture-Recipe Search Using Predictor PretrainingXiaoliang Dai, Alvin Wan, Peizhao Zhang, Bichen Wu et al.CVPR 2021
- Fast and Practical Neural Architecture SearchJiequan Cui, Pengguang Chen, Ruiyu Li, Shu Liu et al.ICCV 2019 · 69 citations
- GreedyNAS: Towards Fast One-Shot NAS With Greedy SupernetShan You, Tao Huang, Mingmin Yang, Fei Wang et al.CVPR 2020
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object DetectionLewei Yao, Hang Xu, Wei Zhang, Xiaodan Liang et al.AAAI 2020 · 83 citations
