GradSign: Model Performance Inference with Theoretical Insights
Zhihao Zhang, Zhihao Jia
Abstract
A key challenge in neural architecture search (NAS) is quickly inferring the predictive performance of a broad spectrum of neural networks to discover statistically accurate and computationally efficient ones. We refer to this task as model performance inference (MPI). The current practice for efficient MPI is gradient-based methods that leverage the gradients of a network at initialization to infer its performance. However, existing gradient-based methods rely only on heuristic metrics and lack the necessary theoretical foundations to consolidate their designs. We propose GradSign, an accurate, simple, and flexible metric for model performance inference with theoretical insights. A key idea behind GradSign is a quantity Ψ to analyze the sample-wise optimization landscape of different networks. Theoretically, we show that Ψ is an upper bound for both the training and true population losses of a neural network under reasonable assumptions. However, it is computationally prohibitive to directly calculate Ψ for modern neural networks. To address this challenge, we design GradSign, an accurate and simple approximation of Ψ using the gradients of a network evaluated at a random initialization state. Evaluation on seven NAS benchmarks across three training datasets shows that GradSign generalizes well to real-world neural networks and consistently outperforms state-of-the-art gradient-based methods for MPI evaluated by Spearman's ρ and Kendall's Tau. Additionally, we have integrated GradSign into four existing NAS algorithms and show that the GradSign-assisted NAS algorithms outperform their vanilla counterparts by improving the accuracies of best-discovered networks by up to 0.3%, 1.1%, and 1.0% on three real-world tasks. Code is available at https://github.com/cmu-catalyst/GradSign
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5e2f9ee-679c-44ae-acf7-0a2cbab41758Cited by top-tier papers19
- Unifying and Boosting Gradient-Based Training-Free Neural Architecture SearchYao Shu, Zhongxiang Dai, Zhaoxuan Wu, Bryan Kian Hsiang LowNeurIPS 2022 · 41 citations
- Training-Free Quantum Architecture SearchZhimin He, Maijie Deng, Shenggen Zheng, Lvzhou Li et al.AAAI 2024 · 39 citations
- MeCo: Zero-Shot NAS with One Data and Single Forward Pass via Minimum Eigenvalue of CorrelationTangyu Jiang, Haodi Wang, Rongfang BieNeurIPS 2023 · 32 citations
- ZiCo: Zero-shot NAS via inverse Coefficient of Variation on GradientsGuihong Li, Yuedong Yang, Kartikeya Bhardwaj, Radu MarculescuICLR 2023 · 19 citations
- Towards Theoretically Inspired Neural Initialization OptimizationYibo Yang, Hong Wang, Haobo Yuan, Zhouchen LinNeurIPS 2022 · 15 citations
Builds on15
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
Related papers
- Neural Architecture Search with Representation Mutual InformationXiawu Zheng, Xiang Fei, Lei Zhang, Chenglin Wu et al.CVPR 2022 · 17 citations
- GP-NAS: Gaussian Process Based Neural Architecture SearchZhihang Li, Teng Xi, Jiankang Deng, Gang Zhang et al.CVPR 2020
- How Powerful are Performance Predictors in Neural Architecture Search?Colin White, Arber Zela, Robin Ru, Yang Liu et al.NeurIPS 2021 · 168 citations
- Loss Functions for Predictor-Based Neural Architecture SearchHan Ji, Yuqi Feng, Jiahao Fan, Yanan SunICCV 2025
- Neural Architecture Search in A Proxy Validation Loss LandscapeYanxi Li, Minjing Dong, Yunhe Wang, Chang XuICML 2020 · 32 citations
