Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NAS
Zihao Sun, Yu Sun, Longxing Yang, Shun Lu, Jilin Mei, Wenxiao Zhao, Yu Hu
Abstract
Neural Architecture Search (NAS) aims to automatically find optimal neural network architectures in an efficient way. Zero-Shot NAS is a promising technique that leverages proxies to predict the accuracy of candidate architectures without any training. However, we have observed that most existing proxies do not consistently perform well across different search spaces, and are less concerned with generalization. Recently, the gradient signal-to-noise ratio (GSNR) was shown to be correlated with neural network generalization performance. In this paper, we not only explicitly give the probability that larger GSNR at network initialization can ensure better generalization, but also theoretically prove that GSNR can ensure better convergence. Then we design the ξ-based gradient signal-to-noise ratio (ξ-GSNR) as a Zero-Shot NAS proxy to predict the network accuracy at initialization. Extensive experiments in different search spaces demonstrate that ξ-GSNR provides superior ranking consistency compared to previous proxies. Moreover, ξ-GSNR-based Zero-Shot NAS also achieves outstanding performance when directly searching for the optimal architecture in various search spaces and datasets. The source code is available at https://github.com/Sunzh1996/Xi-GSNR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae0706b5-7ad4-46d2-bc9d-fce70138c09fCited by top-tier papers9
- GradPower: Powering Gradients for Faster Language Model Pre-TrainingJinbo Wang, Mingze Wang, Jiaqi Zhang, Wei Wang et al.ICML 2026 · 4 citations
- Frozen Language Models Are Gradient Coherence Rectifiers in Vision TransformersLichen Bai, Zixuan Xiong, Hai Lin, Guangwei Xu et al.AAAI 2025 · 4 citations
- L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision TransformersSofia Casarin, Sergio Escalera, Oswald LanzCVPR 2025
- DPaI: Differentiable Pruning at Initialization with Node-Path Balance PrincipleLichuan Xiang, Quan Nguyen-Tri, Lan-Cuong Nguyen, Hoang Pham et al.ICLR 2025
- On the Difficulty of Learning a Meta-network for Training Data SelectionZilin Du, Junqi Zhao, Albert Boyang LiICML 2026
Builds on32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
Related papers
- ZiCo: Zero-shot NAS via inverse Coefficient of Variation on GradientsGuihong Li, Yuedong Yang, Kartikeya Bhardwaj, Radu MarculescuICLR 2023 · 19 citations
- Zen-NAS: A Zero-Shot NAS for High-Performance Image RecognitionMing Lin, Pichao Wang, Zhenhong Sun, Hesen Chen et al.ICCV 2021 · 164 citations
- EZNAS: Evolving Zero-Cost Proxies For Neural Architecture ScoringYash Akhauri, Juan Pablo Muñoz, Nilesh Jain, Ravi IyerNeurIPS 2022 · 15 citations
- MeCo: Zero-Shot NAS with One Data and Single Forward Pass via Minimum Eigenvalue of CorrelationTangyu Jiang, Haodi Wang, Rongfang BieNeurIPS 2023 · 32 citations
- NEAR: A Training-Free Pre-Estimator of Machine Learning Model PerformanceRaphael T. Husistein, Markus Reiher, Marco EckhoffICLR 2025
