On the Global Convergence of Training Deep Linear ResNets
Difan Zou, Philip M. Long, Quanquan Gu
Abstract
We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training -hidden-layer linear residual networks (ResNets). We prove that for training deep residual networks with certain linear transformations at input and output layers, which are fixed throughout training, both GD and SGD with zero initialization on all hidden weights can converge to the global minimum of the training loss. Moreover, when specializing to appropriate Gaussian random linear transformations, GD and SGD provably optimize wide enough deep linear ResNets. Compared with the global convergence result of GD for training standard deep linear networks (Du & Hu 2019), our condition on the neural network width is sharper by a factor of , where denotes the condition number of the covariance matrix of the training data. We further propose a modified identity input and output transformations, and show that a -wide neural network is sufficient to guarantee the global convergence of GD/SGD, where are the input and output dimensions respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78760942-33c1-41b0-a232-c8bb7a10e467Cited by top-tier papers17
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 87 citations
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 51 citations
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 47 citations
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear NetworkJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICML 2021 · 26 citations
Builds on1
Related papers
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 43 citations
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 14 citations
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 82 citations
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu et al.ICML 2020 · 85 citations
