FP-NAS: Fast Probabilistic Neural Architecture Search
Zhicheng Yan, Xiaoliang Dai, Peizhao Zhang, Yuandong Tian, Bichen Wu, Matt Feiszli
摘要
Differential Neural Architecture Search (NAS) requires all layer choices to be held in memory simultaneously; this limits the size of both search space and final architecture. In contrast, Probabilistic NAS, such as PARSEC, learns a distribution over high-performing architectures, and uses only as much memory as needed to train a single model. Nevertheless, it needs to sample many architectures, making it computationally expensive for searching in an extensive space. To solve these problems, we propose a sampling method adaptive to the distribution entropy, drawing more samples to encourage explorations at the beginning, and reducing samples as learning proceeds. Furthermore, to search fast in the multi-variate space, we propose a coarse-to-fine strategy by using a factorized distribution at the beginning which can reduce the number of architecture parameters by over an order of magnitude. We call this method Fast Probabilistic NAS (FP-NAS). Compared with PARSEC, it can sample 64% fewer architectures and search 2.1× faster. Compared with FBNetV2, FP-NAS is 1.9× -3.5× faster, and the searched models outperform FBNetV2 models on ImageNet. FP-NAS allows us to expand the giant FBNetV2 space to be wider (i.e. larger channel choices) and deeper (i.e. more blocks), while adding Split-Attention block and enabling the search over the number of splits. When searching a model of size 0.4G FLOPS, FP-NAS is 132× faster than EfficientNet, and the searched FP-NAS-L0 model outperforms EfficientNet-B0 by 0.7% accuracy. Without using any architecture surrogate or scaling tricks, we directly search large models up to 1.0G FLOPS. Our FP-NAS-L2 model with simple distillation outperforms BigNAS-XL with advanced inplace distillation by 0.7% accuracy using similar FLOPS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Evolving Search Space for Neural Architecture SearchYuanzheng Ci, Chen Lin, Ming Sun, Boyu Chen 等ICCV 2021 · 被引用 48 次
- Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained AnalysisWuyang Chen, Wei Huang, Xinyu Gong, Boris Hanin 等NeurIPS 2022 · 被引用 9 次
- Multi-agent Architecture Search via Agentic SupernetGuibin Zhang, Luyang Niu, Junfeng Fang, Kun Wang 等ICML 2025
它引用的顶会 Paper5
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
- AtomNAS: Fine-Grained End-to-End Neural Architecture SearchJieru Mei, Yingwei Li, Xiaochen Lian, Xiaojie Jin 等ICLR 2020 · 被引用 110 次
- FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel DimensionsAlvin Wan, Xiaoliang Dai, Peizhao Zhang, Zijian He 等CVPR 2020
- SpineNet: Learning Scale-Permuted Backbone for Recognition and LocalizationXianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi 等CVPR 2020
相关 Paper
- AttentiveNAS: Improving Neural Architecture Search via Attentive SamplingDilin Wang, Meng Li, Chengyue Gong, Vikas ChandraCVPR 2021
- Fast and Practical Neural Architecture SearchJiequan Cui, Pengguang Chen, Ruiyu Li, Shu Liu 等ICCV 2019 · 被引用 69 次
- FBNetV3: Joint Architecture-Recipe Search Using Predictor PretrainingXiaoliang Dai, Alvin Wan, Peizhao Zhang, Bichen Wu 等CVPR 2021
- AlphaNet: Improved Training of Supernets with Alpha-DivergenceDilin Wang, Chengyue Gong, Meng Li, Qiang Liu 等ICML 2021 · 被引用 52 次
- Block-Wisely Supervised Neural Architecture Search With Knowledge DistillationChanglin Li, Jiefeng Peng, Liuchun Yuan, Guangrun Wang 等CVPR 2020
