PA&DA: Jointly Sampling PAth and DAta for Consistent NAS
Shun Lu, Yu Hu, Longxing Yang, Zihao Sun, Jilin Mei, Jianchao Tan, Chengru Song
Abstract
Based on the weight-sharing mechanism, one-shot NAS methods train a supernet and then inherit the pre-trained weights to evaluate sub-models, largely reducing the search cost. However, several works have pointed out that the shared weights suffer from different gradient descent directions during training. And we further find that large gradient variance occurs during supernet training, which degrades the supernet ranking consistency. To mitigate this issue, we propose to explicitly minimize the gradient variance of the supernet training by jointly optimizing the sampling distributions of PAth and DAta (PA&DA). We theoretically derive the relationship between the gradient variance and the sampling distributions, and reveal that the optimal sampling probability is proportional to the normalized gradient norm of path and training data. Hence, we use the normalized gradient norm as the importance indicator for path and training data, and adopt an importance sampling strategy for the supernet training. Our method only requires negligible computation cost for optimizing the sampling distributions of path and data, but achieves lower gradient variance during supernet training and better generalization performance for the supernet, resulting in a more consistent NAS. We conduct comprehensive comparisons with other improved approaches in various search spaces. Results show that our method surpasses others with more reliable ranking performance and higher accuracy of searched architectures, showing the effectiveness of our method. Code is available at https://github.com/ShunLu91/PA-DA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36e1442f-db7c-46c8-8f4f-63ea4de6be2bCited by top-tier papers6
- Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear FunctionsYoungmin Oh, Hyunju Lee, Bumsub HamAAAI 2025 · 4 citations
- Boosting Order-Preserving and Transferability for Neural Architecture Search: A Joint Architecture Refined Search and Fine-Tuning ApproachBeichen Zhang, Xiaoxing Wang, Xiaohan Qin, Junchi YanCVPR 2024 · 2 citations
- CARL: Causality-Guided Architecture Representation Learning for an Interpretable Performance PredictorHan Ji, Yuqi Feng, Jiahao Fan, Yanan SunICCV 2025 · 1 citation
- TRNAS: A Training-Free Robust Neural Architecture SearchYeming Yang, Qingling Zhu, Jianping Luo, Ka-Chun Wong et al.ICCV 2025 · 1 citation
- Subnet-Aware Dynamic Supernet Training for Neural Architecture SearchJeimin Jeon, Youngmin Oh, Junghyup Lee, Donghyeon Baek et al.CVPR 2025
Builds on17
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 362 citations
- Rethinking Architecture Selection in Differentiable NASRuochen Wang, Minhao Cheng, Xiangning Chen, Xiaocheng Tang et al.ICLR 2021 · 213 citations
- Few-Shot Neural Architecture SearchYiyang Zhao, Linnan Wang, Yuandong Tian, Rodrigo Fonseca et al.ICML 2021 · 100 citations
Related papers
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang et al.CVPR 2022 · 9 citations
- ShiftNAS: Improving One-shot NAS via Probability ShiftMingyang Zhang, Xinyi Yu, Haodong Zhao, Linlin OuICCV 2023 · 9 citations
- GreedyNAS: Towards Fast One-Shot NAS With Greedy SupernetShan You, Tao Huang, Mingmin Yang, Fei Wang et al.CVPR 2020
- SUMNAS: Supernet with Unbiased Meta-Features for Neural Architecture SearchHyeonmin Ha, Ji-Hoon Kim, Semin Park, Byung-Gon ChunICLR 2022 · 5 citations
- Generalizing Few-Shot NAS with Gradient MatchingShoukang Hu, Ruochen Wang, Lanqing Hong, Zhenguo Li et al.ICLR 2022 · 29 citations
