Celo2: Towards Learned Optimization Free Lunch
Abhinav Moudgil, Boris Knyazev, Eugene Belilovsky
摘要
Learned optimizers are powerful alternatives to hand-designed update rules like Adam, yet they have seen limited practical adoption since they often fail to meta-generalize beyond their training distribution and incur high meta-training cost. For instance, prior work, VeLO, scaled meta-training to 4,000 TPU months (10 GPT-3 compute) to meta-train a general-purpose optimizer but it failed to generalize beyond 600M parameters tasks. In this work, we present a surprising finding: by crafting a simple normalized optimizer architecture and augmenting meta-training, it becomes feasible to meta-train a performant general-purpose learned update rule on a tiny fraction of VeLO compute, 4.5 GPU hours to be precise. Our learned update rule scales stably to a billion-scale pretraining task (GPT-3 XL 1.3B) which is six orders of magnitude larger than its meta-training distribution. Furthermore, it shows strong performance across diverse out-of-distribution tasks and is compatible with modern optimization harness that includes orthogonalization, distinct update rules for input-output and hidden weights, and decoupled weight decay. In all, this work paves the way for practically applicable learnable optimization algorithms, unlocking exploration of richer meta-training and data curation recipes to further improve performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz 等NeurIPS 2022 · 被引用 134 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson 等NeurIPS 2025 · 被引用 46 次
相关 Paper
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon 等ICLR 2026 · 被引用 10 次
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang 等ICLR 2023
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong 等ICML 2024 · 被引用 7 次
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun 等NeurIPS 2021 · 被引用 27 次
- Symbolic Learning to Optimize: Towards Interpretability and ScalabilityWenqing Zheng, Tianlong Chen, Ting-Kuei Hu, Zhangyang WangICLR 2022 · 被引用 21 次
