Generalization Properties of NAS under Activation and Skip Connection Search
Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
摘要
Neural Architecture Search (NAS) has fostered the automatic discovery of state-of-the-art neural architectures. Despite the progress achieved with NAS, so far there is little attention to theoretical guarantees on NAS. In this work, we study the generalization properties of NAS under a unifying framework enabling (deep) layer skip connection search and activation function search. To this end, we derive the lower (and upper) bounds of the minimum eigenvalue of the Neural Tangent Kernel (NTK) under the (in)finite-width regime using a certain search space including mixed activation functions, fully connected, and residual neural networks. We use the minimum eigenvalue to establish generalization error bounds of NAS in the stochastic gradient descent training. Importantly, we theoretically and experimentally show how the derived results can guide NAS to select the top-performing architectures, even in the case without training, leading to a train-free algorithm based on our theory. Accordingly, our numerical validation shed light on the design of computationally efficient methods for NAS. Our analysis is non-trivial due to the coupling of various architectures and activation functions under the unifying framework and has its own interest in providing the lower bound of the minimum eigenvalue of NTK in deep learning theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MeCo: Zero-Shot NAS with One Data and Single Forward Pass via Minimum Eigenvalue of CorrelationTangyu Jiang, Haodi Wang, Rongfang BieNeurIPS 2023 · 被引用 32 次
- On the Convergence of Encoder-only Shallow TransformersYongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2023 · 被引用 17 次
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello 等ICML 2023 · 被引用 12 次
- Robust NAS under adversarial training: benchmark, theory, and beyondYongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel, Grigorios Chrysos 等ICLR 2024 · 被引用 10 次
- MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture SearchYuming Zhang, Jun-Wei Hsieh, Xin Li, Ming-Ching Chang 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper19
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen 等ICLR 2020 · 被引用 691 次
- Neural Architecture Search without TrainingJoe Mellor, Jack Turner, Amos Storkey, Elliot J. CrowleyICML 2021 · 被引用 477 次
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 被引用 327 次
- On Network Design Spaces for Visual RecognitionIlija Radosavovic, Justin Johnson, Saining Xie, Wan-Yen Lo 等ICCV 2019 · 被引用 148 次
相关 Paper
- Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired PerspectiveWuyang Chen, Xinyu Gong, Zhangyang WangICLR 2021 · 被引用 51 次
- Generalization Guarantees for Neural Architecture Search with Train-Validation SplitSamet Oymak, Mingchen Li, Mahdi SoltanolkotabiICML 2021 · 被引用 20 次
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 被引用 98 次
- NASI: Label- and Data-agnostic Neural Architecture Search at InitializationYao Shu, Shaofeng Cai, Zhongxiang Dai, Beng Chin Ooi 等ICLR 2022 · 被引用 51 次
- Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot LearningHaoxiang Wang, Yite Wang, Ruoyu Sun, Bo LiCVPR 2022 · 被引用 37 次
