A Closer Look at Learned Optimization: Stability, Robustness, and Inductive Biases
James Harrison, Luke Metz, Jascha Sohl-Dickstein
Abstract
Learned optimizers -- neural networks that are trained to act as optimizers -- have the potential to dramatically accelerate training of machine learning models. However, even when meta-trained across thousands of tasks at huge computational expense, blackbox learned optimizers often struggle with stability and generalization when applied to tasks unlike those in their meta-training set. In this paper, we use tools from dynamical systems to investigate the inductive biases and stability properties of optimization algorithms, and apply the resulting insights to designing inductive biases for blackbox optimizers. Our investigation begins with a noisy quadratic model, where we characterize conditions in which optimization is stable, in terms of eigenvalues of the training dynamics. We then introduce simple modifications to a learned optimizer's architecture and meta-training procedure which lead to improved stability, and improve the optimizer's inductive bias. We apply the resulting learned optimizer to a variety of neural network training tasks, where it outperforms the current state of the art learned optimizer -- at matched optimizer computational overhead -- with regard to optimization performance and meta-training speed, and is capable of generalization to tasks far different from those it was meta-trained on.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d339b86f-4ea3-4980-94cd-6f8ee5d45dd6Cited by top-tier papers12
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen et al.ICLR 2024 · 57 citations
- Universal Neural FunctionalsAllan Zhou, Chelsea Finn, James HarrisonNeurIPS 2024 · 27 citations
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin et al.ICML 2023 · 18 citations
- Variance-Reduced Gradient Estimation via Noise-Reuse in Online Evolution StrategiesOscar Li, James Harrison, Jascha Sohl-Dickstein, Virginia Smith et al.NeurIPS 2023 · 11 citations
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon et al.ICLR 2026 · 10 citations
Builds on12
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 195 citations
- Continuous Meta-Learning without TasksJames Harrison, Apoorva Sharma, Chelsea Finn, Marco PavoneNeurIPS 2020 · 86 citations
Related papers
- : Unlocking the Performance Ceiling for Pretrained OptimizersMuqi Han, Ruoqi Xing, KAI WU, Xiaoyu Zhang et al.ICML 2026
- Celo2: Towards Learned Optimization Free LunchAbhinav Moudgil, Boris Knyazev, Eugene BelilovskyICLR 2026 · 1 citation
- Guarantees for Tuning the Step Size using a Learning-to-Learn ApproachXiang Wang, Shuai Yuan, Chenwei Wu, Rong GeICML 2021 · 16 citations
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun et al.NeurIPS 2021 · 27 citations
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang et al.ICLR 2023
