Learning Provably Improves the Convergence of Gradient Descent
Qingyu Song, Wei Lin, Hong Xu
Abstract
Learn to Optimize (L2O) trains deep neural network-based solvers for optimization, achieving success in accelerating convex problems and improving non-convex solutions. However, L2O lacks rigorous theoretical backing for its own training convergence, as existing analyses often use unrealistic assumptions-a gap this work highlights empirically. We bridge this gap by proving the training convergence of L2O models that learn Gradient Descent (GD) hyperparameters for quadratic programming, leveraging the Neural Tangent Kernel (NTK) theory. We propose a deterministic initialization strategy to support our theoretical results and promote stable training over extended optimization horizons by mitigating gradient explosion. Our L2O framework demonstrates over 50% better optimality than GD and superior robustness over state-of-the-art L2O methods on synthetic datasets. The code of our method can be found from https://github.com/NetX-lab/MathL2OProof-Official. * This work was done at The Chinese University of Hong Kong (CUHK) when Qingyu was a PhD candidate. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f96a474-1966-42c3-b7ea-d0043f2cfc3bBuilds on9
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 82 citations
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- Hyperparameter Tuning is All You Need for LISTAXiaohan Chen, Jialin Liu, Zhangyang Wang, Wotao YinNeurIPS 2021 · 40 citations
- Safeguarded Learned Convex OptimizationHoward Heaton, Xiaohan Chen, Zhangyang Wang, Wotao YinAAAI 2023 · 33 citations
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin et al.ICML 2023 · 18 citations
Related papers
- Towards Robust Learning to Optimize with Theoretical GuaranteesQingyu Song, Wei Lin, Juncheng Wang, Hong XuCVPR 2024 · 1 citation
- Learning to Optimize Differentiable GamesXuxi Chen, Nelson Vadori, Tianlong Chen, Zhangyang WangICML 2023 · 1 citation
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 15 citations
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 82 citations
- Guarantees for Tuning the Step Size using a Learning-to-Learn ApproachXiang Wang, Shuai Yuan, Chenwei Wu, Rong GeICML 2021 · 16 citations
