Reverse engineering learned optimizers reveals known and novel mechanisms
Niru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun, Jascha Sohl-Dickstein
Abstract
Learned optimizers are parametric algorithms that can themselves be trained to solve optimization problems. In contrast to baseline optimizers (such as momentum or Adam) that use simple update rules derived from theoretical principles, learned optimizers use flexible, high-dimensional, nonlinear parameterizations. Although this can lead to better performance, their inner workings remain a mystery. How is a given learned optimizer able to outperform a well tuned baseline? Has it learned a sophisticated combination of existing optimization techniques, or is it implementing completely new behavior? In this work, we address these questions by careful analysis and visualization of learned optimizers. We study learned optimizers trained from scratch on four disparate tasks, and discover that they have learned interpretable behavior, including: momentum, gradient clipping, learning rate schedules, and learning rate adaptation. Moreover, we show how dynamics and mechanisms inside of learned optimizers orchestrate these computations. Our results help elucidate the previously murky understanding of how learned optimizers work, and establish tools for interpreting future learned optimizers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f8dc3bb-3c2d-4e32-8a9f-deb1158fc52aCited by top-tier papers9
- A Closer Look at Learned Optimization: Stability, Robustness, and Inductive BiasesJames Harrison, Luke Metz, Jascha Sohl-DicksteinNeurIPS 2022 · 41 citations
- Symbolic Learning to Optimize: Towards Interpretability and ScalabilityWenqing Zheng, Tianlong Chen, Ting-Kuei Hu, Zhangyang WangICLR 2022 · 21 citations
- Learn2Hop: Learned Optimization on Rough LandscapesAmil Merchant, Luke Metz, Samuel S. Schoenholz, Ekin D. CubukICML 2021 · 19 citations
- Meta-Learning Bidirectional Update RulesMark Sandler, Max Vladymyrov, Andrey Zhmoginov, Nolan Miller et al.ICML 2021 · 17 citations
- An Operator Theoretic Approach for Analyzing Sequence Neural NetworksIlan Naiman, Omri AzencotAAAI 2023 · 14 citations
Builds on5
- Discovering Symbolic Models from Deep Learning with Inductive BiasesMiles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Rui Xu et al.NeurIPS 2020 · 736 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Reverse-engineering recurrent neural network solutions to a hierarchical inference task for miceRylan Schaeffer, Mikail Khona, Leenoy Meshulam, International Brain Laboratory et al.NeurIPS 2020 · 49 citations
- How recurrent networks implement contextual processing in sentiment analysisNiru Maheswaranathan, David SussilloICML 2020 · 25 citations
Related papers
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon et al.ICLR 2026 · 10 citations
- Celo2: Towards Learned Optimization Free LunchAbhinav Moudgil, Boris Knyazev, Eugene BelilovskyICLR 2026 · 1 citation
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong et al.ICML 2024 · 7 citations
- Towards Constituting Mathematical Structures for Learning to OptimizeJialin Liu, Xiaohan Chen, Zhangyang Wang, Wotao Yin et al.ICML 2023 · 18 citations
- Adaptive Momentum by Momentum for Deep Neural Network TrainingTao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li et al.KDD 2026 · 1 citation
