Prodigy: An Expeditiously Adaptive Parameter-Free Learner
Konstantin Mishchenko, Aaron Defazio
摘要
We consider the problem of estimating the learning rate in adaptive methods, such as AdaGrad and Adam. We propose Prodigy, an algorithm that provably estimates the distance to the solution , which is needed to set the learning rate optimally. At its core, Prodigy is a modification of the D-Adaptation method for learning-rate-free learning. It improves upon the convergence rate of D-Adaptation by a factor of , where is the initial estimate of . We test Prodigy on 12 common logistic-regression benchmark datasets, VGG11 and ResNet-50 training on CIFAR10, ViT training on Imagenet, LSTM training on IWSLT14, DLRM training on Criteo dataset, VarNet on Knee MRI dataset, as well as RoBERTa and GPT transformer training on BookWiki. Our experimental results show that our approach consistently outperforms D-Adaptation and reaches test accuracy values close to that of hand-tuned Adam.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper56
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
- OminiControl: Minimal and Universal Control for Diffusion TransformerZhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue 等ICCV 2025 · 被引用 34 次
- IntrinsiX: High-Quality PBR Generation using Image PriorsPeter Kocsis, Lukas Höllein, Matthias NießnerNeurIPS 2025 · 被引用 20 次
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance GenerationRuihang Xu, Dewei Zhou, Fan Ma, Yi YangICLR 2026 · 被引用 19 次
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 被引用 171 次
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 被引用 117 次
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li 等NeurIPS 2021 · 被引用 111 次
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 被引用 98 次
相关 Paper
- DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent MethodAhmed Khaled, Konstantin Mishchenko, Chi JinNeurIPS 2023 · 被引用 49 次
- Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGradSayantan Choudhury, Nazarii Tupitsa, Nicolas Loizou, Samuel Horváth 等NeurIPS 2024 · 被引用 10 次
- Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuningQitao Tan, Jun Liu, Zheng Zhan, Caiwen Ding 等NeurIPS 2025 · 被引用 19 次
- Gradient descent with generalized Newton's methodZhiqi Bu, Shiyun XuICLR 2025
- Domain-Independent Dominance of Adaptive MethodsPedro Savarese, David McAllester, Sudarshan Babu, Michael MaireCVPR 2021
