Restricted Strong Convexity of Deep Learning Models with Smooth Activations
Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu, Mikhail Belkin
摘要
We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the near initialization'' perspective, we shed considerable new light on the problem. In particular, we make two key technical contributions for such models with $L$ layers, $m$ width, and $σ_0^2$ initialization variance. First, for suitable $σ_0^2$, we establish a $O(\frac{\text{poly}(L)}{\sqrt{m}})$ upper bound on the spectral norm of the Hessian of such models, considerably sharpening prior results. Second, we introduce a new analysis of optimization based on Restricted Strong Convexity (RSC) which holds as long as the squared norm of the average gradient of predictors is $Ω(\frac{\text{poly}(L)}{\sqrt{m}})$ for the square loss. We also present results for more general losses. The RSC based analysis does not need the near initialization" perspective and guarantees geometric convergence for gradient descent (GD). To the best of our knowledge, ours is the first result on establishing geometric convergence of GD based on RSC for deep learning models, thus becoming an alternative sufficient condition for convergence that does not depend on the widely-used Neural Tangent Kernel (NTK). We share preliminary experimental results supporting our theoretical advances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Sketching for Distributed Deep Learning: A Sharper AnalysisMayank Shrivastava, Berivan Isik, Qiaobo Li, Sanmi Koyejo 等NeurIPS 2024 · 被引用 8 次
- Contextual Bandits with Online Neural RegressionRohan Deb, Yikun Ban, Shiliang Zuo, Jingrui He 等ICLR 2024 · 被引用 8 次
- Block Coordinate Descent for Neural Networks Provably Finds Global MinimaShunta AkiyamaNeurIPS 2025 · 被引用 2 次
- Optimization for Neural Operators can Benefit from WidthPedro Cisneros-Velarde, Bhavesh Shrimali, Arindam BanerjeeICML 2025
- Sharper Guarantees for Learning Neural Network Classifiers with Gradient MethodsHossein Taheri, Christos Thrampoulidis, Arya MazumdarICLR 2025
它引用的顶会 Paper8
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
相关 Paper
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 被引用 21 次
- Optimization and Adaptive Generalization of Three layer Neural NetworksKhashayar Gatmiry, Stefanie Jegelka, Jonathan A. KelnerICLR 2022 · 被引用 1 次
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 被引用 44 次
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 被引用 98 次
