Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
Xiang Wang, Shuai Yuan, Chenwei Wu, Rong Ge
摘要
Choosing the right parameters for optimization algorithms is often the key to their success in practice. Solving this problem using a learning-to-learn approach-using meta-gradient descent on a meta-objective based on the trajectory that the optimizer generates-was recently shown to be effective. However, the meta-optimization problem is difficult. In particular, the meta-gradient can often explode/vanish, and the learned optimizer may not have good generalization performance if the meta-objective is not chosen carefully. In this paper we give meta-optimization guarantees for the learning-to-learn approach on a simple problem of tuning the step size for quadratic loss. Our results show that the naïve objective suffers from meta-gradient explosion/vanishing problem. Although there is a way to design the meta-objective so that the meta-gradient remains polynomially bounded, computing the meta-gradient directly using backpropagation leads to numerical issues. We also characterize when it is necessary to compute the meta-objective on a separate validation set to ensure the generalization performance of the learned optimizer. Finally, we verify our results empirically and show that a similar phenomenon appears even for more complicated learned optimizers parametrized by neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 被引用 117 次
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li 等NeurIPS 2021 · 被引用 111 次
- How Important is the Train-Validation Split in Meta-Learning?Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao 等ICML 2021 · 被引用 60 次
- MAML and ANIL Provably Learn RepresentationsLiam Collins, Aryan Mokhtari, Sewoong Oh, Sanjay ShakkottaiICML 2022 · 被引用 38 次
相关 Paper
- MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parametersArsalan Sharifnassab, Saber Salehkaleybar, Richard S. SuttonICML 2025
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon 等ICLR 2026 · 被引用 10 次
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 被引用 66 次
- Optimistic Meta-GradientsSebastian Flennerhag, Tom Zahavy, Brendan O'Donoghue, Hado Philip van Hasselt 等NeurIPS 2023 · 被引用 3 次
- Online Control for Meta-optimizationXinyi Chen, Elad HazanNeurIPS 2023 · 被引用 9 次
