Leveraging the two-timescale regime to demonstrate convergence of neural networks
Pierre Marion, Raphaël Berthier
Abstract
We study the training dynamics of shallow neural networks, in a two-timescale regime in which the stepsizes for the inner layer are much smaller than those for the outer layer. In this regime, we prove convergence of the gradient flow to a global optimum of the non-convex optimization problem in a simple univariate setting. The number of neurons need not be asymptotically large for our result to hold, distinguishing our result from popular recent approaches such as the neural tangent kernel or mean-field regimes. Experimental illustration is provided, showing that the stochastic gradient descent behaves according to our description of the gradient flow and thus converges to a global optimum in the two-timescale regime, but can fail outside of this regime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97d2f87d-a980-4235-a04f-628fa5835e27Cited by top-tier papers13
- Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention LandscapeJuno Kim, Taiji SuzukiICML 2024 · 42 citations
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 15 citations
- Mean-field Analysis on Two-layer Neural Networks from a Kernel PerspectiveShokichi Takakura, Taiji SuzukiICML 2024 · 12 citations
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala et al.NeurIPS 2025 · 10 citations
- Mean-Field Langevin Dynamics for Signed Measures via a Bilevel ApproachGuillaume Wang, Alireza Mousavi-Hosseini, Lénaïc ChizatNeurIPS 2024 · 8 citations
Builds on4
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- AutoLR: Layer-wise Pruning and Auto-tuning of Learning Rates in Fine-tuning of Deep NetworksYoungmin Ro, Jin Young ChoiAAAI 2021 · 44 citations
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 26 citations
- Not All Layers Are Equal: A Layer-Wise Adaptive Approach Toward Large-Scale DNN TrainingYun-Yong Ko, Dongwon Lee, Sang-Wook KimWWW 2022 · 11 citations
Related papers
- On feature learning in neural networks with global convergence guaranteesZhengdao Chen, Eric Vanden-Eijnden, Joan BrunaICLR 2022 · 15 citations
- The Convex Geometry of Backpropagation: Neural Network Gradient Flows Converge to Extreme Points of the Dual Convex ProgramYifei Wang, Mert PilanciICLR 2022 · 12 citations
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 92 citations
- Global Convergence of Three-layer Neural Networks in the Mean Field RegimeHuy Tuan Pham, Phan-Minh NguyenICLR 2021 · 23 citations
- Rethinking Neural Network Learning Rates: A Stackelberg PerspectiveSihan Zeng, Sujay Bhatt, Sumitra GaneshICML 2026
