Generalization Guarantees for Neural Architecture Search with Train-Validation Split
Samet Oymak, Mingchen Li, Mahdi Soltanolkotabi
摘要
Neural Architecture Search (NAS) is a popular method for automatically designing optimized architectures for high-performance deep learning. In this approach, it is common to use bilevel optimization where one optimizes the model weights over the training data (lower-level problem) and various hyperparameters such as the configuration of the architecture over the validation data (upper-level problem). This paper explores the statistical aspects of such problems with train-validation splits. In practice, the lower-level problem is often overparameterized and can easily achieve zero loss. Thus, a-priori it seems impossible to distinguish the right hyperparameters based on training loss alone which motivates a better understanding of the role of train-validation split. To this aim this work establishes the following results: • We show that refined properties of the validation loss such as risk and hyper-gradients are indicative of those of the true test loss. This reveals that the upper-level problem helps select the most generalizable model and prevent overfitting with a near-minimal validation sample size. Importantly, this is established for continuous search spaces which are relevant for popular differentiable search schemes. Extensions to transfer learning are developed in terms of the mismatch between training & validation distributions. • We establish generalization bounds for NAS problems with an emphasis on an activation search problem. When optimized with gradient-descent, we show that the train-validation procedure returns the best (model, architecture) pair even if all architectures can perfectly fit the training data to achieve zero error. • Finally, we highlight rigorous connections between NAS, multiple kernel learning, and low-rank matrix learning. The latter leads to novel algorithmic insights where the solution of the upper problem can be accurately learned via efficient spectral methods to achieve near-minimal risk.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- AutoBalance: Optimized Loss Functions for Imbalanced DataMingchen Li, Xuechen Zhang, Christos Thrampoulidis, Jiasi Chen 等NeurIPS 2021 · 被引用 89 次
- Generalization Properties of NAS under Activation and Skip Connection SearchZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 被引用 23 次
- Robust NAS under adversarial training: benchmark, theory, and beyondYongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel, Grigorios Chrysos 等ICLR 2024 · 被引用 10 次
- On the Privacy Risks of Cell-Based NAS ArchitecturesHai Huang, Zhikun Zhang, Yun Shen, Michael Backes 等CCS 2022 · 被引用 6 次
- CONTRAST: Continual Multi-source Adaptation to Dynamic DistributionsSk Miraj Ahmed, Fahim Faisal Niloy, Xiangyu Chang, Dripta S. Raychaudhuri 等NeurIPS 2024 · 被引用 5 次
它引用的顶会 Paper10
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen 等ICLR 2020 · 被引用 691 次
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 被引用 107 次
- Neural Kernels Without TangentsVaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil 等ICML 2020 · 被引用 93 次
- Compressive sensing with un-trained neural networks: Gradient descent finds a smooth approximationReinhard Heckel, Mahdi SoltanolkotabiICML 2020 · 被引用 91 次
相关 Paper
- Improving Differentiable Neural Architecture Search by Encouraging TransferabilityParth Sheth, Pengtao XieICLR 2023
- Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot LearningHaoxiang Wang, Yite Wang, Ruoyu Sun, Bo LiCVPR 2022 · 被引用 37 次
- Towards Accurate and Robust Architectures via Neural Architecture SearchYuwei Ou, Yuqi Feng, Yanan SunCVPR 2024 · 被引用 8 次
- Analyzing Generalization of Neural Networks through Loss Path KernelsYilan Chen, Wei Huang, Hao Wang, Charlotte Loh 等NeurIPS 2023 · 被引用 4 次
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse CodingYibo Yang, Hongyang Li, Shan You, Fei Wang 等NeurIPS 2020 · 被引用 66 次
