Generalization Guarantees for Neural Architecture Search with Train-Validation Split
Samet Oymak, Mingchen Li, Mahdi Soltanolkotabi
Abstract
Neural Architecture Search (NAS) is a popular method for automatically designing optimized architectures for high-performance deep learning. In this approach, it is common to use bilevel optimization where one optimizes the model weights over the training data (lower-level problem) and various hyperparameters such as the configuration of the architecture over the validation data (upper-level problem). This paper explores the statistical aspects of such problems with train-validation splits. In practice, the lower-level problem is often overparameterized and can easily achieve zero loss. Thus, a-priori it seems impossible to distinguish the right hyperparameters based on training loss alone which motivates a better understanding of the role of train-validation split. To this aim this work establishes the following results: • We show that refined properties of the validation loss such as risk and hyper-gradients are indicative of those of the true test loss. This reveals that the upper-level problem helps select the most generalizable model and prevent overfitting with a near-minimal validation sample size. Importantly, this is established for continuous search spaces which are relevant for popular differentiable search schemes. Extensions to transfer learning are developed in terms of the mismatch between training & validation distributions. • We establish generalization bounds for NAS problems with an emphasis on an activation search problem. When optimized with gradient-descent, we show that the train-validation procedure returns the best (model, architecture) pair even if all architectures can perfectly fit the training data to achieve zero error. • Finally, we highlight rigorous connections between NAS, multiple kernel learning, and low-rank matrix learning. The latter leads to novel algorithmic insights where the solution of the upper problem can be accurately learned via efficient spectral methods to achieve near-minimal risk.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff758fd5-625d-4228-9dd8-9ad72136a688Cited by top-tier papers7
- AutoBalance: Optimized Loss Functions for Imbalanced DataMingchen Li, Xuechen Zhang, Christos Thrampoulidis, Jiasi Chen et al.NeurIPS 2021 · 89 citations
- Generalization Properties of NAS under Activation and Skip Connection SearchZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 23 citations
- Robust NAS under adversarial training: benchmark, theory, and beyondYongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel, Grigorios Chrysos et al.ICLR 2024 · 10 citations
- On the Privacy Risks of Cell-Based NAS ArchitecturesHai Huang, Zhikun Zhang, Yun Shen, Michael Backes et al.CCS 2022 · 6 citations
- CONTRAST: Continual Multi-source Adaptation to Dynamic DistributionsSk Miraj Ahmed, Fahim Faisal Niloy, Xiangyu Chang, Dripta S. Raychaudhuri et al.NeurIPS 2024 · 5 citations
Builds on10
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen et al.ICLR 2020 · 691 citations
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 107 citations
- Neural Kernels Without TangentsVaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil et al.ICML 2020 · 93 citations
- Compressive sensing with un-trained neural networks: Gradient descent finds a smooth approximationReinhard Heckel, Mahdi SoltanolkotabiICML 2020 · 91 citations
Related papers
- Improving Differentiable Neural Architecture Search by Encouraging TransferabilityParth Sheth, Pengtao XieICLR 2023
- Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot LearningHaoxiang Wang, Yite Wang, Ruoyu Sun, Bo LiCVPR 2022 · 37 citations
- Towards Accurate and Robust Architectures via Neural Architecture SearchYuwei Ou, Yuqi Feng, Yanan SunCVPR 2024 · 8 citations
- Analyzing Generalization of Neural Networks through Loss Path KernelsYilan Chen, Wei Huang, Hao Wang, Charlotte Loh et al.NeurIPS 2023 · 4 citations
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse CodingYibo Yang, Hongyang Li, Shan You, Fei Wang et al.NeurIPS 2020 · 66 citations
