DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
Ahmed Khaled, Konstantin Mishchenko, Chi Jin
Abstract
This paper proposes a new easy-to-implement parameter-free gradient-based optimizer: DoWG (Distance over Weighted Gradients). We prove that DoWG is efficient -- matching the convergence rate of optimally tuned gradient descent in convex optimization up to a logarithmic factor without tuning any parameters, and universal -- automatically adapting to both smooth and nonsmooth problems. While popular algorithms following the AdaGrad framework compute a running average of the squared gradients to use for normalization, DoWG maintains a new distance-based weighted version of the running average, which is crucial to achieve the desired properties. To complement our theory, we also show empirically that DoWG trains at the edge of stability, and validate its effectiveness on practical machine learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05b6d48c-2dfa-4868-843e-a667f72e0374Cited by top-tier papers17
- Adaptive Proximal Gradient Method for Convex OptimizationYura Malitsky, Konstantin MishchenkoNeurIPS 2024 · 80 citations
- Universality of AdaGrad Stepsizes for Stochastic Optimization: Inexact Oracle, Acceleration and Variance ReductionAnton Rodomanov, Xiaowen Jiang, Sebastian U. StichNeurIPS 2024 · 14 citations
- SGD with Adaptive Preconditioning: Unified Analysis and Momentum AccelerationDmitry KovalevICLR 2026 · 13 citations
- Tuning-Free Stochastic OptimizationAhmed Khaled, Chi JinICML 2024 · 13 citations
- Parameter-free Clipped Gradient Descent Meets PolyakYuki Takezawa, Han Bao, Ryoma Sato, Kenta Niwa et al.NeurIPS 2024 · 11 citations
Builds on11
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 171 citations
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 132 citations
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
Related papers
- Accelerated Distance-adaptive Methods for Hölder Smooth and Convex OptimizationYijin Ren, Haifeng Xu, Qi DengNeurIPS 2025
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
- DADA: Dual Averaging with Distance AdaptationMohammad Moshtaghifar, Anton Rodomanov, Daniil Vankov, Sebastian U StichICLR 2026 · 4 citations
- A Parameter-Free and Near-Optimal Zeroth-Order Algorithm for Stochastic Convex OptimizationKunjie Ren, Luo LuoICML 2025
- A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean DescentShuo Xie, Tianhao Wang, Beining Wu, Zhiyuan LiICLR 2026 · 7 citations
