Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
Xiang Wang, Shuai Yuan, Chenwei Wu, Rong Ge
Abstract
Choosing the right parameters for optimization algorithms is often the key to their success in practice. Solving this problem using a learning-to-learn approach-using meta-gradient descent on a meta-objective based on the trajectory that the optimizer generates-was recently shown to be effective. However, the meta-optimization problem is difficult. In particular, the meta-gradient can often explode/vanish, and the learned optimizer may not have good generalization performance if the meta-objective is not chosen carefully. In this paper we give meta-optimization guarantees for the learning-to-learn approach on a simple problem of tuning the step size for quadratic loss. Our results show that the naïve objective suffers from meta-gradient explosion/vanishing problem. Although there is a way to design the meta-objective so that the meta-gradient remains polynomially bounded, computing the meta-gradient directly using backpropagation leads to numerical issues. We also characterize when it is necessary to compute the meta-objective on a separate validation set to ensure the generalization performance of the learned optimizer. Finally, we verify our results empirically and show that a similar phenomenon appears even for more complicated learned optimizers parametrized by neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f84b225-9a75-46ad-aec0-ab4e2fb53b44Cited by top-tier papers11
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li et al.NeurIPS 2021 · 111 citations
- How Important is the Train-Validation Split in Meta-Learning?Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao et al.ICML 2021 · 60 citations
- MAML and ANIL Provably Learn RepresentationsLiam Collins, Aryan Mokhtari, Sewoong Oh, Sanjay ShakkottaiICML 2022 · 38 citations
Related papers
- MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parametersArsalan Sharifnassab, Saber Salehkaleybar, Richard S. SuttonICML 2025
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon et al.ICLR 2026 · 10 citations
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 66 citations
- Optimistic Meta-GradientsSebastian Flennerhag, Tom Zahavy, Brendan O'Donoghue, Hado Philip van Hasselt et al.NeurIPS 2023 · 3 citations
- Online Control for Meta-optimizationXinyi Chen, Elad HazanNeurIPS 2023 · 9 citations
