ZiCo: Zero-shot NAS via inverse Coefficient of Variation on Gradients
Guihong Li, Yuedong Yang, Kartikeya Bhardwaj, Radu Marculescu
摘要
Neural Architecture Search (NAS) is widely used to automatically obtain the neural network with the best performance among a large number of candidate architectures. To reduce the search time, zero-shot NAS aims at designing training-free proxies that can predict the test performance of a given architecture. However, as shown recently, none of the zero-shot proxies proposed to date can actually work consistently better than a naive proxy, namely, the number of network parameters (#Params). To improve this state of affairs, as the main theoretical contribution, we first reveal how some specific gradient properties across different samples impact the convergence rate and generalization capacity of neural networks. Based on this theoretical analysis, we propose a new zero-shot proxy, ZiCo, the first proxy that works consistently better than #Params. We demonstrate that ZiCo works better than State-Of-The-Art (SOTA) proxies on several popular NAS-Benchmarks (NASBench101, NATSBench-SSS/TSS, TransNASBench-101) for multiple applications (e.g., image classification/reconstruction and pixel-level prediction). Finally, we demonstrate that the optimal architectures found via ZiCo are as competitive as the ones found by one-shot and multi-shot NAS methods, but with much less search time. For example, ZiCo-based NAS can find optimal architectures with 78.1%, 79.4%, and 80.4% test accuracy under inference budgets of 450M, 600M, and 1000M FLOPs, respectively, on ImageNet within 0.4 GPU days. Our code is available at https://github.com/SLDGroup/ZiCo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- MeCo: Zero-Shot NAS with One Data and Single Forward Pass via Minimum Eigenvalue of CorrelationTangyu Jiang, Haodi Wang, Rongfang BieNeurIPS 2023 · 被引用 32 次
- SWAP-NAS: Sample-Wise Activation Patterns for Ultra-fast NASYameng Peng, Andy Song, Haytham M. Fayek, Vic Ciesielski 等ICLR 2024 · 被引用 22 次
- ParZC: Parametric Zero-Cost Proxies for Efficient NASPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu 等AAAI 2025 · 被引用 14 次
- Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NASZihao Sun, Yu Sun, Longxing Yang, Shun Lu 等ICCV 2023 · 被引用 10 次
- On skip connections and normalisation layers in deep optimisationLachlan E. MacDonald, Jack Valmadre, Hemanth Saratchandran, Simon LuceyNeurIPS 2023 · 被引用 8 次
它引用的顶会 Paper38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
相关 Paper
- Zen-NAS: A Zero-Shot NAS for High-Performance Image RecognitionMing Lin, Pichao Wang, Zhenhong Sun, Hesen Chen 等ICCV 2021 · 被引用 164 次
- SiGeo: Sub-One-Shot NAS via Geometry of Loss LandscapeHua Zheng, Kuang-Hung Liu, Igor Fedorov, Xin Zhang 等KDD 2024 · 被引用 1 次
- Zero-Cost Proxies for Lightweight NASMohamed S. Abdelfattah, Abhinav Mehrotra, Lukasz Dudziak, Nicholas Donald LaneICLR 2021 · 被引用 65 次
- Extensible and Efficient Proxy for Neural Architecture SearchYuhong Li, Jiajie Li, Cong Hao, Pan Li 等ICCV 2023 · 被引用 8 次
- AZ-NAS: Assembling Zero-Cost Proxies for Network Architecture SearchJunghyup Lee, Bumsub HamCVPR 2024
