Generalization Properties of NAS under Activation and Skip Connection Search
Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
Abstract
Neural Architecture Search (NAS) has fostered the automatic discovery of state-of-the-art neural architectures. Despite the progress achieved with NAS, so far there is little attention to theoretical guarantees on NAS. In this work, we study the generalization properties of NAS under a unifying framework enabling (deep) layer skip connection search and activation function search. To this end, we derive the lower (and upper) bounds of the minimum eigenvalue of the Neural Tangent Kernel (NTK) under the (in)finite-width regime using a certain search space including mixed activation functions, fully connected, and residual neural networks. We use the minimum eigenvalue to establish generalization error bounds of NAS in the stochastic gradient descent training. Importantly, we theoretically and experimentally show how the derived results can guide NAS to select the top-performing architectures, even in the case without training, leading to a train-free algorithm based on our theory. Accordingly, our numerical validation shed light on the design of computationally efficient methods for NAS. Our analysis is non-trivial due to the coupling of various architectures and activation functions under the unifying framework and has its own interest in providing the lower bound of the minimum eigenvalue of NTK in deep learning theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f72a2e7b-2ee3-45a7-ab4e-6a536e11c3a6Cited by top-tier papers7
- MeCo: Zero-Shot NAS with One Data and Single Forward Pass via Minimum Eigenvalue of CorrelationTangyu Jiang, Haodi Wang, Rongfang BieNeurIPS 2023 · 32 citations
- On the Convergence of Encoder-only Shallow TransformersYongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2023 · 17 citations
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello et al.ICML 2023 · 12 citations
- Robust NAS under adversarial training: benchmark, theory, and beyondYongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel, Grigorios Chrysos et al.ICLR 2024 · 10 citations
- MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture SearchYuming Zhang, Jun-Wei Hsieh, Xin Li, Ming-Ching Chang et al.NeurIPS 2024 · 4 citations
Builds on19
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen et al.ICLR 2020 · 691 citations
- Neural Architecture Search without TrainingJoe Mellor, Jack Turner, Amos Storkey, Elliot J. CrowleyICML 2021 · 477 citations
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 327 citations
- On Network Design Spaces for Visual RecognitionIlija Radosavovic, Justin Johnson, Saining Xie, Wan-Yen Lo et al.ICCV 2019 · 148 citations
Related papers
- Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired PerspectiveWuyang Chen, Xinyu Gong, Zhangyang WangICLR 2021 · 51 citations
- Generalization Guarantees for Neural Architecture Search with Train-Validation SplitSamet Oymak, Mingchen Li, Mahdi SoltanolkotabiICML 2021 · 20 citations
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 98 citations
- NASI: Label- and Data-agnostic Neural Architecture Search at InitializationYao Shu, Shaofeng Cai, Zhongxiang Dai, Beng Chin Ooi et al.ICLR 2022 · 51 citations
- Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot LearningHaoxiang Wang, Yite Wang, Ruoyu Sun, Bo LiCVPR 2022 · 37 citations
