Demystify Hyperparameters for Stochastic Optimization with Transferable Representations
Jianhui Sun, Mengdi Huai, Kishlay Jha, Aidong Zhang
Abstract
This paper studies the convergence and generalization of a large class of Stochastic Gradient Descent (SGD) momentum schemes, in both learning from scratch and transferring representations with fine-tuning. Momentum-based acceleration of SGD is the default optimizer for many deep learning models. However, there is a lack of general convergence guarantees for many existing momentum variants in conjunction withstochastic gradient. It is also unclear how the momentum methods may affect thegeneralization error. In this paper, we give a unified analysis of several popular optimizers, e.g., Polyak's heavy ball momentum and Nesterov's accelerated gradient. Our contribution is threefold. First, we give a unified convergence guarantee for a large class of momentum variants in thestochastic setting. Notably, our results cover both convex and nonconvex objectives. Second, we prove a generalization bound for neural networks trained by momentum variants. We analyze how hyperparameters affect the generalization bound and consequently propose guidelines on how to tune these hyperparameters in various momentum schemes to generalize well. We provide extensive empirical evidence to our proposed guidelines. Third, this study fills the vacancy of a formal analysis of fine-tuning in literature. To our best knowledge, our work is the first systematic generalizability analysis on momentum methods that cover both learning from scratch and fine-tuning. Our codes are available https://github.com/jsycsjh/Demystify-Hyperparameters-for-Stochastic-Optimization-with-Transferable-Representations .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c80d7a51-dbc0-4c66-987c-c1f18c620f9aCited by top-tier papers4
- Faster Adaptive Federated LearningXidong Wu, Feihu Huang, Zhengmian Hu, Heng HuangAAAI 2023 · 99 citations
- Decentralized Riemannian Algorithm for Nonconvex Minimax ProblemsXidong Wu, Zhengmian Hu, Heng HuangAAAI 2023 · 15 citations
- Serverless Federated AUPRC Optimization for Multi-Party Collaborative Imbalanced Data MiningXidong Wu, Zhengmian Hu, Jian Pei, Heng HuangKDD 2023 · 13 citations
- Enhance Diffusion to Improve Robust GeneralizationJianhui Sun, Sanchit Sinha, Aidong ZhangKDD 2023 · 1 citation
Builds on5
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
- On the Convergence of Nesterov's Accelerated Gradient Method in Stochastic SettingsMahmoud Assran, Mike RabbatICML 2020 · 71 citations
- Correlation Networks for Extreme Multi-label Text ClassificationGuangxu Xun, Kishlay Jha, Jianhui Sun, Aidong ZhangKDD 2020 · 60 citations
- Malicious Attacks against Deep Reinforcement Learning InterpretationsMengdi Huai, Jianhui Sun, Renqin Cai, Liuyi Yao et al.KDD 2020 · 27 citations
- A Stagewise Hyperparameter Scheduler to Improve GeneralizationJianhui Sun, Ying Yang, Guangxu Xun, Aidong ZhangKDD 2021 · 8 citations
Related papers
- Adaptive Momentum by Momentum for Deep Neural Network TrainingTao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li et al.KDD 2026 · 1 citation
- The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball MethodsWei Tao, Sheng Long, Gaowei Wu, Qing TaoICLR 2021 · 17 citations
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu et al.ICLR 2024 · 14 citations
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic ModelsCourtney Paquette, Elliot PaquetteNeurIPS 2021 · 20 citations
- High Probability Bounds for Non-Convex Stochastic Optimization with MomentumShaojie Li, Pengwei Tang, Bowei Zhu, Yong LiuICLR 2026 · 100 citations
