Restricted Strong Convexity of Deep Learning Models with Smooth Activations
Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu, Mikhail Belkin
Abstract
We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the near initialization'' perspective, we shed considerable new light on the problem. In particular, we make two key technical contributions for such models with $L$ layers, $m$ width, and $σ_0^2$ initialization variance. First, for suitable $σ_0^2$, we establish a $O(\frac{\text{poly}(L)}{\sqrt{m}})$ upper bound on the spectral norm of the Hessian of such models, considerably sharpening prior results. Second, we introduce a new analysis of optimization based on Restricted Strong Convexity (RSC) which holds as long as the squared norm of the average gradient of predictors is $Ω(\frac{\text{poly}(L)}{\sqrt{m}})$ for the square loss. We also present results for more general losses. The RSC based analysis does not need the near initialization" perspective and guarantees geometric convergence for gradient descent (GD). To the best of our knowledge, ours is the first result on establishing geometric convergence of GD based on RSC for deep learning models, thus becoming an alternative sufficient condition for convergence that does not depend on the widely-used Neural Tangent Kernel (NTK). We share preliminary experimental results supporting our theoretical advances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0569730f-9965-42a2-bbf6-9fa4f9c80c93Cited by top-tier papers6
- Sketching for Distributed Deep Learning: A Sharper AnalysisMayank Shrivastava, Berivan Isik, Qiaobo Li, Sanmi Koyejo et al.NeurIPS 2024 · 8 citations
- Contextual Bandits with Online Neural RegressionRohan Deb, Yikun Ban, Shiliang Zuo, Jingrui He et al.ICLR 2024 · 8 citations
- Block Coordinate Descent for Neural Networks Provably Finds Global MinimaShunta AkiyamaNeurIPS 2025 · 2 citations
- Optimization for Neural Operators can Benefit from WidthPedro Cisneros-Velarde, Bhavesh Shrimali, Arindam BanerjeeICML 2025
- Sharper Guarantees for Learning Neural Network Classifiers with Gradient MethodsHossein Taheri, Christos Thrampoulidis, Arya MazumdarICLR 2025
Builds on8
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 260 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
Related papers
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 21 citations
- Optimization and Adaptive Generalization of Three layer Neural NetworksKhashayar Gatmiry, Stefanie Jegelka, Jonathan A. KelnerICLR 2022 · 1 citation
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 44 citations
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 98 citations
