Better Parameter-Free Stochastic Optimization with ODE Updates for Coin-Betting
Keyi Chen, John Langford, Francesco Orabona
Abstract
Parameter-free stochastic gradient descent (PFSGD) algorithms do not require setting learning rates while achieving optimal theoretical performance. In practical applications, however, there remains an empirical gap between tuned stochastic gradient descent (SGD) and PFSGD. In this paper, we close the empirical gap with a new parameter-free algorithm based on continuous-time Coin-Betting on truncated models. The new update is derived through the solution of an Ordinary Differential Equation (ODE) and solved in a closed form. We show empirically that this new parameter-free algorithm outperforms algorithms with the ``best default'' learning rates and almost matches the performance of finely tuned baselines without anything to tune.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92ad445f-3874-47e1-8540-77a7ae448fcaCited by top-tier papers9
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
- PDE-Based Optimal Strategy for Unconstrained Online LearningZhiyu Zhang, Ashok Cutkosky, Ioannis Ch. PaschalidisICML 2022 · 31 citations
- Auditing Fairness by BettingBen Chugg, Santiago Cortes-Gomez, Bryan Wilder, Aaditya RamdasNeurIPS 2023 · 29 citations
- How Free is Parameter-Free Stochastic Optimization?Amit Attia, Tomer KorenICML 2024 · 11 citations
- Adaptive Variance Reduction for Stochastic Optimization under Weaker AssumptionsWei Jiang, Sifan Yang, Yibo Wang, Lijun ZhangNeurIPS 2024 · 11 citations
Related papers
- Coin Sampling: Gradient-Based Bayesian Inference without Learning RatesLouis Sharrock, Christopher NemethICML 2023 · 10 citations
- Parameter-free Clipped Gradient Descent Meets PolyakYuki Takezawa, Han Bao, Ryoma Sato, Kenta Niwa et al.NeurIPS 2024 · 11 citations
- Tuning-Free Stochastic OptimizationAhmed Khaled, Chi JinICML 2024 · 13 citations
- Learning Rate Free Bayesian Inference in Constrained DomainsLouis Sharrock, Lester Mackey, Christopher NemethNeurIPS 2023 · 3 citations
- DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent MethodAhmed Khaled, Konstantin Mishchenko, Chi JinNeurIPS 2023 · 49 citations
