On global convergence of ResNets: From finite to infinite width using linear parameterization
Raphaël Barboni, Gabriel Peyré, François-Xavier Vialard
摘要
Overparameterization is a key factor in the absence of convexity to explain global convergence of gradient descent (GD) for neural networks. Beside the well studied lazy regime, infinite width (mean field) analysis has been developed for shallow networks, using convex optimization techniques. To bridge the gap between the lazy and mean field regimes, we study Residual Networks (ResNets) in which the residual block has linear parameterization while still being nonlinear. Such ResNets admit both infinite depth and width limits, encoding residual blocks in a Reproducing Kernel Hilbert Space (RKHS). In this limit, we prove a local Polyak-Lojasiewicz inequality. Thus, every critical point is a global minimizer and a local convergence result of GD holds, retrieving the lazy regime. In contrast with other mean-field studies, it applies to both parametric and non-parametric cases under an expressivity condition on the residuals. Our analysis leads to a practical and quantified recipe: starting from a universal RKHS, Random Fourier Features are applied to obtain a finite dimensional parameterization satisfying with highprobability our expressivity condition. Related works and contributions Recently, several works have addressed the problem of proving convergence of (stochastic) GD in the training of NNs. If the convergence properties of GD are well understood for NNs that are linear w.r.t. input [24, 7, 64] , it is not the case for non-linear NNs. In [34, 33, 17] , the authors focus on the training of "shallow" two layers fully connected NNs and establish convergence of GD in an
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 被引用 24 次
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He 等NeurIPS 2024 · 被引用 10 次
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos 等ICLR 2024 · 被引用 2 次
它引用的顶会 Paper9
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 被引用 44 次
- Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient DescentSpencer Frei, Quanquan GuNeurIPS 2021 · 被引用 30 次
相关 Paper
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari 等NeurIPS 2021 · 被引用 35 次
- Gradient Descent in Neural Networks as Sequential Learning in Reproducing Kernel Banach SpaceAlistair Shilton, Sunil Gupta, Santu Rana, Svetha VenkateshICML 2023 · 被引用 3 次
- A Recipe for Global Convergence Guarantee in Deep Neural NetworksKenji Kawaguchi, Qingyun SunAAAI 2021 · 被引用 14 次
- On feature learning in neural networks with global convergence guaranteesZhengdao Chen, Eric Vanden-Eijnden, Joan BrunaICLR 2022 · 被引用 15 次
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
