Learning Provably Improves the Convergence of Gradient Descent
Qingyu Song, Wei Lin, Hong Xu
摘要
Learn to Optimize (L2O) trains deep neural network-based solvers for optimization, achieving success in accelerating convex problems and improving non-convex solutions. However, L2O lacks rigorous theoretical backing for its own training convergence, as existing analyses often use unrealistic assumptions-a gap this work highlights empirically. We bridge this gap by proving the training convergence of L2O models that learn Gradient Descent (GD) hyperparameters for quadratic programming, leveraging the Neural Tangent Kernel (NTK) theory. We propose a deterministic initialization strategy to support our theoretical results and promote stable training over extended optimization horizons by mitigating gradient explosion. Our L2O framework demonstrates over 50% better optimality than GD and superior robustness over state-of-the-art L2O methods on synthetic datasets. The code of our method can be found from https://github.com/NetX-lab/MathL2OProof-Official. * This work was done at The Chinese University of Hong Kong (CUHK) when Qingyu was a PhD candidate. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 被引用 82 次
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- Hyperparameter Tuning is All You Need for LISTAXiaohan Chen, Jialin Liu, Zhangyang Wang, Wotao YinNeurIPS 2021 · 被引用 40 次
- Safeguarded Learned Convex OptimizationHoward Heaton, Xiaohan Chen, Zhangyang Wang, Wotao YinAAAI 2023 · 被引用 33 次
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin 等ICML 2023 · 被引用 18 次
相关 Paper
- Towards Robust Learning to Optimize with Theoretical GuaranteesQingyu Song, Wei Lin, Juncheng Wang, Hong XuCVPR 2024 · 被引用 1 次
- Learning to Optimize Differentiable GamesXuxi Chen, Nelson Vadori, Tianlong Chen, Zhangyang WangICML 2023 · 被引用 1 次
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 被引用 15 次
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 被引用 82 次
- Guarantees for Tuning the Step Size using a Learning-to-Learn ApproachXiang Wang, Shuai Yuan, Chenwei Wu, Rong GeICML 2021 · 被引用 16 次
