Reverse engineering learned optimizers reveals known and novel mechanisms
Niru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun, Jascha Sohl-Dickstein
摘要
Learned optimizers are parametric algorithms that can themselves be trained to solve optimization problems. In contrast to baseline optimizers (such as momentum or Adam) that use simple update rules derived from theoretical principles, learned optimizers use flexible, high-dimensional, nonlinear parameterizations. Although this can lead to better performance, their inner workings remain a mystery. How is a given learned optimizer able to outperform a well tuned baseline? Has it learned a sophisticated combination of existing optimization techniques, or is it implementing completely new behavior? In this work, we address these questions by careful analysis and visualization of learned optimizers. We study learned optimizers trained from scratch on four disparate tasks, and discover that they have learned interpretable behavior, including: momentum, gradient clipping, learning rate schedules, and learning rate adaptation. Moreover, we show how dynamics and mechanisms inside of learned optimizers orchestrate these computations. Our results help elucidate the previously murky understanding of how learned optimizers work, and establish tools for interpreting future learned optimizers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Closer Look at Learned Optimization: Stability, Robustness, and Inductive BiasesJames Harrison, Luke Metz, Jascha Sohl-DicksteinNeurIPS 2022 · 被引用 41 次
- Symbolic Learning to Optimize: Towards Interpretability and ScalabilityWenqing Zheng, Tianlong Chen, Ting-Kuei Hu, Zhangyang WangICLR 2022 · 被引用 21 次
- Learn2Hop: Learned Optimization on Rough LandscapesAmil Merchant, Luke Metz, Samuel S. Schoenholz, Ekin D. CubukICML 2021 · 被引用 19 次
- Meta-Learning Bidirectional Update RulesMark Sandler, Max Vladymyrov, Andrey Zhmoginov, Nolan Miller 等ICML 2021 · 被引用 17 次
- An Operator Theoretic Approach for Analyzing Sequence Neural NetworksIlan Naiman, Omri AzencotAAAI 2023 · 被引用 14 次
它引用的顶会 Paper5
- Discovering Symbolic Models from Deep Learning with Inductive BiasesMiles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Rui Xu 等NeurIPS 2020 · 被引用 736 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Reverse-engineering recurrent neural network solutions to a hierarchical inference task for miceRylan Schaeffer, Mikail Khona, Leenoy Meshulam, International Brain Laboratory 等NeurIPS 2020 · 被引用 49 次
- How recurrent networks implement contextual processing in sentiment analysisNiru Maheswaranathan, David SussilloICML 2020 · 被引用 25 次
相关 Paper
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon 等ICLR 2026 · 被引用 10 次
- Celo2: Towards Learned Optimization Free LunchAbhinav Moudgil, Boris Knyazev, Eugene BelilovskyICLR 2026 · 被引用 1 次
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong 等ICML 2024 · 被引用 7 次
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin 等ICML 2023 · 被引用 18 次
- Adaptive Momentum by Momentum for Deep Neural Network TrainingTao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li 等KDD 2026 · 被引用 1 次
