Celo2: Towards Learned Optimization Free Lunch
Abhinav Moudgil, Boris Knyazev, Eugene Belilovsky
Abstract
Learned optimizers are powerful alternatives to hand-designed update rules like Adam, yet they have seen limited practical adoption since they often fail to meta-generalize beyond their training distribution and incur high meta-training cost. For instance, prior work, VeLO, scaled meta-training to 4,000 TPU months (10 GPT-3 compute) to meta-train a general-purpose optimizer but it failed to generalize beyond 600M parameters tasks. In this work, we present a surprising finding: by crafting a simple normalized optimizer architecture and augmenting meta-training, it becomes feasible to meta-train a performant general-purpose learned update rule on a tiny fraction of VeLO compute, 4.5 GPU hours to be precise. Our learned update rule scales stably to a billion-scale pretraining task (GPT-3 XL 1.3B) which is six orders of magnitude larger than its meta-training distribution. Furthermore, it shows strong performance across diverse out-of-distribution tasks and is compatible with modern optimization harness that includes orthogonalization, distinct update rules for input-output and hidden weights, and decoupled weight decay. In all, this work paves the way for practically applicable learnable optimization algorithms, unlocking exploration of richer meta-training and data curation recipes to further improve performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2f2024e-8801-45bf-af11-3c8ecce8e90fBuilds on7
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz et al.NeurIPS 2022 · 134 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson et al.NeurIPS 2025 · 46 citations
Related papers
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon et al.ICLR 2026 · 10 citations
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang et al.ICLR 2023
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong et al.ICML 2024 · 7 citations
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun et al.NeurIPS 2021 · 27 citations
- Symbolic Learning to Optimize: Towards Interpretability and ScalabilityWenqing Zheng, Tianlong Chen, Ting-Kuei Hu, Zhangyang WangICLR 2022 · 21 citations
