On the Global Convergence of Training Deep Linear ResNets
Difan Zou, Philip M. Long, Quanquan Gu
摘要
We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training -hidden-layer linear residual networks (ResNets). We prove that for training deep residual networks with certain linear transformations at input and output layers, which are fixed throughout training, both GD and SGD with zero initialization on all hidden weights can converge to the global minimum of the training loss. Moreover, when specializing to appropriate Gaussian random linear transformations, GD and SGD provably optimize wide enough deep linear ResNets. Compared with the global convergence result of GD for training standard deep linear networks (Du & Hu 2019), our condition on the neural network width is sharper by a factor of , where denotes the condition number of the covariance matrix of the training data. We further propose a modified identity input and output transformations, and show that a -wide neural network is sufficient to guarantee the global convergence of GD/SGD, where are the input and output dimensions respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 被引用 87 次
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 被引用 51 次
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 被引用 47 次
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear NetworkJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICML 2021 · 被引用 26 次
它引用的顶会 Paper1
相关 Paper
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 被引用 43 次
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 被引用 14 次
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 被引用 82 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
