A Closer Look at Learned Optimization: Stability, Robustness, and Inductive Biases
James Harrison, Luke Metz, Jascha Sohl-Dickstein
摘要
Learned optimizers -- neural networks that are trained to act as optimizers -- have the potential to dramatically accelerate training of machine learning models. However, even when meta-trained across thousands of tasks at huge computational expense, blackbox learned optimizers often struggle with stability and generalization when applied to tasks unlike those in their meta-training set. In this paper, we use tools from dynamical systems to investigate the inductive biases and stability properties of optimization algorithms, and apply the resulting insights to designing inductive biases for blackbox optimizers. Our investigation begins with a noisy quadratic model, where we characterize conditions in which optimization is stable, in terms of eigenvalues of the training dynamics. We then introduce simple modifications to a learned optimizer's architecture and meta-training procedure which lead to improved stability, and improve the optimizer's inductive bias. We apply the resulting learned optimizer to a variety of neural network training tasks, where it outperforms the current state of the art learned optimizer -- at matched optimizer computational overhead -- with regard to optimization performance and meta-training speed, and is capable of generalization to tasks far different from those it was meta-trained on.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen 等ICLR 2024 · 被引用 57 次
- Universal Neural FunctionalsAllan Zhou, Chelsea Finn, James HarrisonNeurIPS 2024 · 被引用 27 次
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin 等ICML 2023 · 被引用 18 次
- Variance-Reduced Gradient Estimation via Noise-Reuse in Online Evolution StrategiesOscar Li, James Harrison, Jascha Sohl-Dickstein, Virginia Smith 等NeurIPS 2023 · 被引用 11 次
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper12
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- Continuous Meta-Learning without TasksJames Harrison, Apoorva Sharma, Chelsea Finn, Marco PavoneNeurIPS 2020 · 被引用 86 次
相关 Paper
- : Unlocking the Performance Ceiling for Pretrained OptimizersMuqi Han, Ruoqi Xing, KAI WU, Xiaoyu Zhang 等ICML 2026
- Celo2: Towards Learned Optimization Free LunchAbhinav Moudgil, Boris Knyazev, Eugene BelilovskyICLR 2026 · 被引用 1 次
- Guarantees for Tuning the Step Size using a Learning-to-Learn ApproachXiang Wang, Shuai Yuan, Chenwei Wu, Rong GeICML 2021 · 被引用 16 次
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun 等NeurIPS 2021 · 被引用 27 次
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang 等ICLR 2023
