A Stagewise Hyperparameter Scheduler to Improve Generalization
Jianhui Sun, Ying Yang, Guangxu Xun, Aidong Zhang
摘要
Stochastic gradient descent (SGD) augmented with various momentum variants (e.g. heavy ball momentum (SHB) and Nesterov's accelerated gradient (NAG)) has been the default optimizer for many learning tasks. Tuning the optimizer's hyperparameters is arguably the most time-consuming part of model training. Many new momentum variants, despite their empirical advantage over classical SHB/NAG, introduce even more hyperparameters to tune. Automating the tedious and error-prone tuning is essential for AutoML. This paper focuses on how to efficiently tune a large class of multistage momentum variants to improve generalization. We use the general formulation of quasi-hyperbolic momentum (QHM) and extend "constant and drop'', the widespread learning rate α scheduler where α is set large initially and then dropped every few epochs, to other hyperparameters (e.g. batch size b, momentum parameter β, instant discount factor ν). Multistage QHM is a unified framework which covers a large family of momentum variants as its special cases (e.g. vanilla SGD/SHB/NAG). Existing works mainly focus on scheduling α's decay, while multistage QHM allows additional varying hyperparameters such as b, β, and ν, and demonstrates better generalization ability than only tuning α. Our tuning strategies have rigorous justifications rather than a blind trial-and-error. We theoretically prove why our tuning strategies could improve generalization. We also show the convergence of multistage QHM for general nonconvex objective functions. Our strategies simplify the tuning process and beat competitive optimizers in test accuracy empirically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 被引用 5 次
- HyperTime: Hyperparameter Optimization for Combating Temporal Distribution ShiftsShaokun Zhang, Yiran Wu, Zhonghua Zheng, Qingyun Wu 等ACM MM 2024 · 被引用 2 次
- Enhance Diffusion to Improve Robust GeneralizationJianhui Sun, Sanchit Sinha, Aidong ZhangKDD 2023 · 被引用 1 次
它引用的顶会 Paper3
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 被引用 328 次
- DeepMV: Multi-View Deep Learning for Device-Free Human Activity RecognitionHongfei Xue, Wenjun Jiang, Chenglin Miao, Fenglong Ma 等UbiComp 2020 · 被引用 64 次
- Correlation Networks for Extreme Multi-label Text ClassificationGuangxu Xun, Kishlay Jha, Jianhui Sun, Aidong ZhangKDD 2020 · 被引用 60 次
相关 Paper
- Adaptive Momentum by Momentum for Deep Neural Network TrainingTao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li 等KDD 2026 · 被引用 1 次
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic ModelsCourtney Paquette, Elliot PaquetteNeurIPS 2021 · 被引用 20 次
- Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient NoiseRui Pan, Yuxing Liu, Xiaoyu Wang, Tong ZhangICLR 2024 · 被引用 10 次
- Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical PerformanceDimitris Oikonomou, Nicolas LoizouICLR 2025
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 被引用 32 次
