NAS evaluation is frustratingly hard
Antoine Yang, Pedro M. Esperança, Fabio Maria Carlucci
摘要
Neural Architecture Search (NAS) is an exciting new field which promises to be as much as a game-changer as Convolutional Neural Networks were in 2012. Despite many great works leading to substantial improvements on a variety of tasks, comparison between different methods is still very much an open issue. While most algorithms are tested on the same datasets, there is no shared experimental protocol followed by all. As such, and due to the under-use of ablation studies, there is a lack of clarity regarding why certain methods are more effective than others. Our first contribution is a benchmark of 8 NAS methods on 5 datasets. To overcome the hurdle of comparing methods with different search spaces, we propose using a method's relative improvement over the randomly sampled average architecture, which effectively removes advantages arising from expertly engineered search spaces or training protocols. Surprisingly, we find that many NAS techniques struggle to significantly beat the average architecture baseline. We perform further experiments with the commonly used DARTS search space in order to understand the contribution of each component in the NAS pipeline. These experiments highlight that: (i) the use of tricks in the evaluation protocol has a predominant impact on the reported performance of architectures; (ii) the cell-based search space has a very narrow accuracy range, such that the seed has a considerable impact on architecture rankings; (iii) the hand-designed macrostructure (cells) is more important than the searched micro-structure (operations); and (iv) the depth-gap is a real phenomenon, evidenced by the change in rankings between 8 and 20 cell architectures. To conclude, we suggest best practices, that we hope will prove useful for the community and help mitigate current NAS pitfalls, e.g. difficulties in reproducibility and comparison of search methods. The code used is available at https://github.com/antoyang/NAS-Benchmark .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- BANANAS: Bayesian Optimization with Neural Architectures for Neural Architecture SearchColin White, Willie Neiswanger, Yash SavaniAAAI 2021 · 被引用 401 次
- AutoML-Zero: Evolving Machine Learning Algorithms From ScratchEsteban Real, Chen Liang, David R. So, Quoc V. LeICML 2020 · 被引用 265 次
- How Powerful are Performance Predictors in Neural Architecture Search?Colin White, Arber Zela, Robin Ru, Yang Liu 等NeurIPS 2021 · 被引用 168 次
- NAS-Bench-1Shot1: Benchmarking and Dissecting One-shot Neural Architecture SearchArber Zela, Julien Siems, Frank HutterICLR 2020 · 被引用 156 次
- Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONASHan Shi, Renjie Pi, Hang Xu, Zhenguo Li 等NeurIPS 2020 · 被引用 148 次
它引用的顶会 Paper2
相关 Paper
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate ImportanceHongyi He, Longjun Liu, Haonan Zhang, Nanning ZhengAAAI 2024 · 被引用 21 次
- Shapley-NAS: Discovering Operation Contribution for Neural Architecture SearchHan Xiao, Ziwei Wang, Zheng Zhu, Jie Zhou 等CVPR 2022 · 被引用 59 次
- Rethink DARTS Search Space and Renovate a New BenchmarkJiuling Zhang, Zhiming DingICML 2023 · 被引用 3 次
- AGNAS: Attention-Guided Micro and Macro-Architecture SearchZihao Sun, Yu Hu, Shun Lu, Longxing Yang 等ICML 2022 · 被引用 16 次
