Learning-Rate-Free Learning by D-Adaptation
Aaron Defazio, Konstantin Mishchenko
Abstract
D-Adaptation is an approach to automatically setting the learning rate which asymptotically achieves the optimal rate of convergence for minimizing convex Lipschitz functions, with no back-tracking or line searches, and no additional function value or gradient evaluations per step. Our approach is the first hyper-parameter free method for this class without additional multiplicative log factors in the convergence rate. We present extensive experiments for SGD and Adam variants of our method, where the method automatically matches hand-tuned learning rates across more than a dozen diverse machine learning problems, including large-scale vision and language problems. An open-source implementation is available 1 . * The work was prepared while K. Mishchenko was at
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8990e9fc-b985-48de-a5d8-e692ccbde5b0Cited by top-tier papers49
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- The Road Less ScheduledAaron Defazio, Xingyu Yang, Ahmed Khaled, Konstantin Mishchenko et al.NeurIPS 2024 · 208 citations
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett et al.ICLR 2024 · 162 citations
- Prodigy: An Expeditiously Adaptive Parameter-Free LearnerKonstantin Mishchenko, Aaron DefazioICML 2024 · 131 citations
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 66 citations
- PDE-Based Optimal Strategy for Unconstrained Online LearningZhiyu Zhang, Ashok Cutkosky, Ioannis Ch. PaschalidisICML 2022 · 31 citations
- Guarantees for Tuning the Step Size using a Learning-to-Learn ApproachXiang Wang, Shuai Yuan, Chenwei Wu, Rong GeICML 2021 · 16 citations
Related papers
- Tuning-Free Stochastic OptimizationAhmed Khaled, Chi JinICML 2024 · 13 citations
- QLABGrad: A Hyperparameter-Free and Convergence-Guaranteed Scheme for Deep LearningMinghan Fu, Fang-Xiang WuAAAI 2024 · 12 citations
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order InformationMajid Jahani, Sergey Rusakov, Zheng Shi, Peter Richtárik et al.ICLR 2022 · 31 citations
- Mechanic: A Learning Rate TunerAshok Cutkosky, Aaron Defazio, Harsh MehtaNeurIPS 2023 · 27 citations
- DADA: Dual Averaging with Distance AdaptationMohammad Moshtaghifar, Anton Rodomanov, Daniil Vankov, Sebastian U StichICLR 2026 · 4 citations
