Acceleration via Fractal Learning Rate Schedules
Naman Agarwal, Surbhi Goel, Cyril Zhang
Abstract
In practical applications of iterative first-order optimization, the learning rate schedule remains notoriously difficult to understand and expensive to tune. We demonstrate the presence of these subtleties even in the innocuous case when the objective is a convex quadratic. We reinterpret an iterative algorithm from the numerical analysis literature as what we call the Chebyshev learning rate schedule for accelerating vanilla gradient descent, and show that the problem of mitigating instability leads to a fractal ordering of step sizes. We provide some experiments to challenge conventional beliefs about stable learning rates in deep learning: the fractal schedule enables training to converge with locally unstable updates which make negative progress on the objective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b60ba03-ddfc-48eb-b45c-82e456702accCited by top-tier papers6
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 77 citations
- Fractal Structure and Generalization Properties of Stochastic Optimization AlgorithmsAlexander Camuto, George Deligiannidis, Murat A. Erdogdu, Mert Gürbüzbalaban et al.NeurIPS 2021 · 34 citations
- The Curse of Unrolling: Rate of Differentiating Through OptimizationDamien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian PedregosaNeurIPS 2022 · 20 citations
- Chaotic Regularization and Heavy-Tailed Limits for Deterministic Gradient DescentSoon Hoe Lim, Yijun Wan, Umut SimsekliNeurIPS 2022 · 14 citations
- Block Acceleration Without Momentum: On Optimal Stepsizes of Block Gradient Descent for Least-SquaresLiangzu Peng, Wotao YinICML 2024 · 3 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 93 citations
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 25 citations
Related papers
- Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian SpectrumsRui Pan, Haishan Ye, Tong ZhangICLR 2022 · 20 citations
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 44 citations
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 15 citations
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning RateZhiqi Bu, Shiyun Xu, Jialin MaoICLR 2026 · 4 citations
- Algorithmic Instabilities of Accelerated Gradient DescentAmit Attia, Tomer KorenNeurIPS 2021 · 21 citations
