Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained Analysis
Wuyang Chen, Wei Huang, Xinyu Gong, Boris Hanin, Zhangyang Wang
Abstract
Advanced deep neural networks (DNNs), designed by either human or AutoML algorithms, are growing increasingly complex. Diverse operations are connected by complicated connectivity patterns, e.g., various types of skip connections. Those topological compositions are empirically effective and observed to smooth the loss landscape and facilitate the gradient flow in general. However, it remains elusive to derive any principled understanding of their effects on the DNN capacity or trainability, and to understand why or in which aspect one specific connectivity pattern is better than another. In this work, we theoretically characterize the impact of connectivity patterns on the convergence of DNNs under gradient descent training in fine granularity. By analyzing a wide network's Neural Network Gaussian Process (NNGP), we are able to depict how the spectrum of an NNGP kernel propagates through a particular connectivity pattern, and how that affects the bound of convergence rates. As one practical implication of our results, we show that by a simple filtration on "unpromising" connectivity patterns, we can trim down the number of models to evaluate, and significantly accelerate the large-scale neural architecture search without any overhead. Code is available at: https://github.com/VITA-Group/architecture_convergence .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0524e1b1-7936-4d48-aefc-745a86d8122cCited by top-tier papers1
Ask how each one uses itBuilds on24
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 725 citations
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen et al.ICLR 2020 · 691 citations
- Neural Architecture Search without TrainingJoe Mellor, Jack Turner, Amos Storkey, Elliot J. CrowleyICML 2021 · 477 citations
Related papers
- Understanding Architectures Learnt by Cell-based Neural Architecture SearchYao Shu, Wei Wang, Shaofeng CaiICLR 2020 · 92 citations
- Generalization Properties of NAS under Activation and Skip Connection SearchZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 23 citations
- Principled Architecture-aware Scaling of HyperparametersWuyang Chen, Junru Wu, Zhangyang Wang, Boris HaninICLR 2024 · 3 citations
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 87 citations
- A Kernel Perspective of Skip Connections in Convolutional NetworksDaniel Barzilai, Amnon Geifman, Meirav Galun, Ronen BasriICLR 2023 · 3 citations
