Gradient Descent Averaging and Primal-dual Averaging for Strongly Convex Optimization
Wei Tao, Wei Li, Zhisong Pan, Qing Tao
Abstract
Averaging scheme has attracted extensive attention in deep learning as well as traditional machine learning. It achieves theoretically optimal convergence and also improves the empirical model performance. However, there is still a lack of sufficient convergence analysis for strongly convex optimization. Typically, the convergence about the last iterate of gradient descent methods, which is referred to as individual convergence, fails to attain its optimality due to the existence of logarithmic factor. In order to remove this factor, we first develop gradient descent averaging (GDA), which is a general projection-based dual averaging algorithm in the strongly convex setting. We further present primal-dual averaging for strongly convex cases (SC-PDA), where primal and dual averaging schemes are simultaneously utilized. We prove that GDA yields the optimal convergence rate in terms of output averaging, while SC-PDA derives the optimal individual convergence. Several experiments on SVMs and deep learning models validate the correctness of theoretical analysis and effectiveness of algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd577c55-0c4a-488a-bf81-19f15d3c050eCited by top-tier papers2
- Symmetric Mean-field Langevin Dynamics for Distributional Minimax ProblemsJuno Kim, Kakei Yamamoto, Kazusato Oko, Zhuoran Yang et al.ICLR 2024 · 14 citations
- Convergence of Mean-Field Langevin Stochastic Descent-Ascent for Distributional Minimax OptimizationZhangyi Liu, Feng Liu, Rui Gao, Shuang LiICML 2025
Builds on1
Related papers
- Variance Reduction via Accelerated Dual Averaging for Finite-Sum OptimizationChaobing Song, Yong Jiang, Yi MaNeurIPS 2020 · 25 citations
- The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball MethodsWei Tao, Sheng Long, Gaowei Wu, Qing TaoICLR 2021 · 17 citations
- Variance Reduction via Primal-Dual Accelerated Dual Averaging for Nonsmooth Convex Finite-SumsChaobing Song, Stephen J. Wright, Jelena DiakonikolasICML 2021 · 22 citations
- Obtaining Adjustable Regularization for Free via Iterate AveragingJingfeng Wu, Vladimir Braverman, Lin YangICML 2020 · 2 citations
- Averaged Method of Multipliers for Bi-Level Optimization without Lower-Level Strong ConvexityRisheng Liu, Yaohua Liu, Wei Yao, Shangzhi Zeng et al.ICML 2023 · 37 citations
