Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian Spectrums
Rui Pan, Haishan Ye, Tong Zhang
Abstract
Learning rate schedulers have been widely adopted in training deep neural networks. Despite their practical importance, there is a discrepancy between its practice and its theoretical analysis. For instance, it is not known what schedules of SGD achieve best convergence, even for simple problems such as optimizing quadratic objectives. In this paper, we propose Eigencurve, the first family of learning rate schedules that can achieve minimax optimal convergence rates (up to a constant) for SGD on quadratic objectives when the eigenvalue distribution of the underlying Hessian matrix is skewed. The condition is quite common in practice. Experimental results show that Eigencurve can significantly outperform step decay in image classification tasks on CIFAR-10, especially when the number of epochs is small. Moreover, the theory inspires two simple learning rate schedulers for practical applications that can approximate eigencurve. For some problems, the optimal shape of the proposed schedulers resembles that of cosine decay, which sheds light to the success of cosine decay for such situations. For other situations, the proposed schedulers are superior to cosine decay.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e68b9d05-0d9f-4062-b807-ea4e350b2ac5Cited by top-tier papers9
- Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear RegressionJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu et al.ICML 2022 · 38 citations
- Revisiting the Last-Iterate Convergence of Stochastic Gradient MethodsZijian Liu, Zhengyuan ZhouICLR 2024 · 32 citations
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio et al.NeurIPS 2024 · 25 citations
- Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient NoiseRui Pan, Yuxing Liu, Xiaoyu Wang, Tong ZhangICLR 2024 · 10 citations
- Stepsize anything: A unified learning rate schedule for budgeted-iteration trainingAnda Tang, Yiming Dong, Yutao Zeng, Xun Zhou et al.NeurIPS 2025 · 1 citation
Builds on5
- Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear RegressionJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu et al.ICML 2022 · 38 citations
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and PerformanceXiaoyu Li, Zhenxun Zhuang, Francesco OrabonaICML 2021 · 29 citations
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of SymmetryYossi Arjevani, Michael FieldNeurIPS 2020 · 22 citations
- SPAN: A Stochastic Projected Approximate Newton MethodXunpeng Huang, Xianfeng Liang, Zhengyang Liu, Lei Li et al.AAAI 2020 · 4 citations
- Enhance Curvature Information by Structured Stochastic Quasi-Newton MethodsMinghan Yang, Dong Xu, Hongyu Chen, Zaiwen Wen et al.CVPR 2021
Related papers
- Acceleration via Fractal Learning Rate SchedulesNaman Agarwal, Surbhi Goel, Cyril ZhangICML 2021 · 19 citations
- Stepping on the Edge: Curvature Aware Learning Rate TunersVincent Roulet, Atish Agarwala, Jean-Bastien Grill, Grzegorz Swirszcz et al.NeurIPS 2024 · 9 citations
- On the Convergence of Step Decay Step-Size for Stochastic OptimizationXiaoyu Wang, Sindri Magnússon, Mikael JohanssonNeurIPS 2021 · 33 citations
- Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning RateJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan GuICLR 2021 · 18 citations
- QLABGrad: A Hyperparameter-Free and Convergence-Guaranteed Scheme for Deep LearningMinghan Fu, Fang-Xiang WuAAAI 2024 · 12 citations
