μLO: Compute-Efficient Meta-Generalization of Learned Optimizers
Benjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon, Irina Rish, Eugene Belilovsky
Abstract
Learned optimizers (LOs) have the potential to significantly reduce the wall-clock training time of neural networks. However, they can struggle to optimize unseen tasks (meta-generalize), especially when training networks wider than those seen during meta-training. To address this, we derive the Maximal Update Parametrization (P) for two state-of-the-art learned optimizer architectures and propose a simple meta-training recipe for -parameterized LOs (LOs). Our empirical evaluation demonstrates that LOs meta-trained with our recipe substantially improve meta-generalization to wider unseen tasks when compared to LOs trained under standard parametrization (SP) using the same compute budget. We also empirically observe that LOs exhibit unexpectedly improved meta-generalization to deeper networks ( meta-training) and surprising generalization to much longer training horizons ( meta-training) when compared to SP LOs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85dcf8b9-ba62-43ac-a480-1dfce0456a24Cited by top-tier papers1
Ask how each one uses itBuilds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 132 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
- Don't be lazy: CompleteP enables compute-efficient deep transformersNolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li et al.NeurIPS 2025 · 77 citations
Related papers
- Celo2: Towards Learned Optimization Free LunchAbhinav Moudgil, Boris Knyazev, Eugene BelilovskyICLR 2026 · 1 citation
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang et al.ICLR 2023
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun et al.NeurIPS 2021 · 27 citations
- A Closer Look at Learned Optimization: Stability, Robustness, and Inductive BiasesJames Harrison, Luke Metz, Jascha Sohl-DicksteinNeurIPS 2022 · 41 citations
- : Unlocking the Performance Ceiling for Pretrained OptimizersMuqi Han, Ruoqi Xing, KAI WU, Xiaoyu Zhang et al.ICML 2026
