On global convergence of ResNets: From finite to infinite width using linear parameterization
Raphaël Barboni, Gabriel Peyré, François-Xavier Vialard
Abstract
Overparameterization is a key factor in the absence of convexity to explain global convergence of gradient descent (GD) for neural networks. Beside the well studied lazy regime, infinite width (mean field) analysis has been developed for shallow networks, using convex optimization techniques. To bridge the gap between the lazy and mean field regimes, we study Residual Networks (ResNets) in which the residual block has linear parameterization while still being nonlinear. Such ResNets admit both infinite depth and width limits, encoding residual blocks in a Reproducing Kernel Hilbert Space (RKHS). In this limit, we prove a local Polyak-Lojasiewicz inequality. Thus, every critical point is a global minimizer and a local convergence result of GD holds, retrieving the lazy regime. In contrast with other mean-field studies, it applies to both parametric and non-parametric cases under an expressivity condition on the residuals. Our analysis leads to a practical and quantified recipe: starting from a universal RKHS, Random Fourier Features are applied to obtain a finite dimensional parameterization satisfying with highprobability our expressivity condition. Related works and contributions Recently, several works have addressed the problem of proving convergence of (stochastic) GD in the training of NNs. If the convergence properties of GD are well understood for NNs that are linear w.r.t. input [24, 7, 64] , it is not the case for non-linear NNs. In [34, 33, 17] , the authors focus on the training of "shallow" two layers fully connected NNs and establish convergence of GD in an
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 24 citations
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He et al.NeurIPS 2024 · 10 citations
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos et al.ICLR 2024 · 2 citations
Builds on9
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu et al.ICML 2020 · 85 citations
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 44 citations
- Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient DescentSpencer Frei, Quanquan GuNeurIPS 2021 · 30 citations
Related papers
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari et al.NeurIPS 2021 · 35 citations
- Gradient Descent in Neural Networks as Sequential Learning in Reproducing Kernel Banach SpaceAlistair Shilton, Sunil Gupta, Santu Rana, Svetha VenkateshICML 2023 · 3 citations
- A Recipe for Global Convergence Guarantee in Deep Neural NetworksKenji Kawaguchi, Qingyun SunAAAI 2021 · 14 citations
- On feature learning in neural networks with global convergence guaranteesZhengdao Chen, Eric Vanden-Eijnden, Joan BrunaICLR 2022 · 15 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
