Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local Smoothness
Kevin Scaman, Cédric Malherbe, Ludovic Dos Santos
Abstract
Training over-parameterized neural networks involves the empirical minimization of highly nonconvex objective functions. Recently, a large body of works provided theoretical evidence that, despite this non-convexity, properly initialized over-parameterized networks can converge to a zero training loss through the introduction of the Polyak-Łojasiewicz condition. However, these analyses are restricted to quadratic losses such as mean square error, and tend to indicate fast exponential convergence rates that are seldom observed in practice. In this work, we propose to extend these results by analyzing stochastic gradient descent under more generic Łojasiewicz conditions that are applicable to any convex loss function, thus extending the current theory to a larger panel of losses commonly used in practice such as cross-entropy. Moreover, our analysis provides high-probability bounds on the approximation error under sub-Gaussian gradient noise and only requires the local smoothness of the objective function, thus making it applicable to deep neural networks in realistic settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Sharp Analysis of Stochastic Optimization under Global Kurdyka-Lojasiewicz InequalityIlyas Fatkhullin, Jalal Etesami, Niao He, Negar KiyavashNeurIPS 2022 · 34 citations
- How to Make the Gradients Small Privately: Improved Rates for Differentially Private Non-Convex OptimizationAndrew Lowy, Jonathan R. Ullman, Stephen J. WrightICML 2024 · 11 citations
- Optimization for Amortized Inverse ProblemsTianci Liu, Tong Yang, Quan Zhang, Qi LeiICML 2023 · 7 citations
- Convergence beyond the over-parameterized regime using Rayleigh quotientsDavid A. R. Robin, Kevin Scaman, Marc LelargeNeurIPS 2022 · 7 citations
- Policy Gradient Methods Converge Globally in Imperfect-Information Extensive-Form GamesFivos Kalogiannis, Gabriele FarinaNeurIPS 2025 · 2 citations
Builds on5
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang et al.ICML 2020 · 122 citations
- Robustness Analysis of Non-Convex Stochastic Gradient Descent using Biased ExpectationsKevin Scaman, Cédric MalherbeNeurIPS 2020 · 37 citations
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?Zixiang Chen, Yuan Cao, Difan Zou, Quanquan GuICLR 2021 · 29 citations
Related papers
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari et al.NeurIPS 2021 · 35 citations
- Loss Landscape Characterization of Neural Networks without Over-ParametrizationRustem Islamov, Niccolò Ajroldi, Antonio Orvieto, Aurélien LucchiNeurIPS 2024 · 14 citations
- Proxy Convexity: A Unified Framework for the Analysis of Neural Networks Trained by Gradient DescentSpencer Frei, Quanquan GuNeurIPS 2021 · 30 citations
- Sharpness-Aware Minimization: General Analysis and Improved RatesDimitris Oikonomou, Nicolas LoizouICLR 2025
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
