Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance Reduction
Radu-Alexandru Dragomir, Mathieu Even, Hadrien Hendrikx
摘要
We study the problem of minimizing a relatively-smooth convex function using stochastic Bregman gradient methods. We first prove the convergence of Bregman Stochastic Gradient Descent (BSGD) to a region that depends on the noise (magnitude of the gradients) at the optimum. In particular, BSGD with a constant step-size converges to the exact minimizer when this noise is zero (interpolation setting, in which the data is fit perfectly). Otherwise, when the objective has a finite sum structure, we show that variance reduction can be used to counter the effect of noise. In particular, fast convergence to the exact minimizer can be obtained under additional regularity assumptions on the Bregman reference function. We illustrate the effectiveness of our approach on two key applications of relative smoothness: tomographic reconstruction with Poisson noise and statistical preconditioning for distributed optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityMathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas FlammarionNeurIPS 2023 · 被引用 42 次
- Stochastic Gradient Descent under Markovian Sampling SchemesMathieu EvenICML 2023 · 被引用 41 次
- On Sample Optimality in Personalized Collaborative and Federated LearningMathieu Even, Laurent Massoulié, Kevin ScamanNeurIPS 2022 · 被引用 24 次
- Stochastic Distributed Optimization under Average Second-order Similarity: Algorithms and AnalysisDachao Lin, Yuze Han, Haishan Ye, Zhihua ZhangNeurIPS 2023 · 被引用 17 次
- Two Losses Are Better Than One: Faster Optimization Using a Cheaper ProxyBlake E. Woodworth, Konstantin Mishchenko, Francis R. BachICML 2023 · 被引用 9 次
它引用的顶会 Paper3
- Dual-Free Stochastic Decentralized Optimization with Variance ReductionHadrien Hendrikx, Francis R. Bach, Laurent MassouliéNeurIPS 2020 · 被引用 29 次
- Regret Bounds without Lipschitz Continuity: Online Learning with Relative-Lipschitz LossesYihan Zhou, Victor S. Portella, Mark Schmidt, Nicholas J. A. HarveyNeurIPS 2020 · 被引用 25 次
- Online and stochastic optimization beyond Lipschitz continuity: A Riemannian approachKimon Antonakopoulos, Elena Veronica Belmega, Panayotis MertikopoulosICLR 2020 · 被引用 20 次
相关 Paper
- Adaptive First-Order Methods Revisited: Convex Minimization without Lipschitz RequirementsKimon Antonakopoulos, Panayotis MertikopoulosNeurIPS 2021 · 被引用 13 次
- Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient DescentSharan Vaswani, Benjamin Dubois-Taine, Reza BabanezhadICML 2022
- Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying RegularizationSebastian Kassing, Simon Weissmann, Leif DöringNeurIPS 2025 · 被引用 7 次
- Fast Stochastic Composite Minimization and an Accelerated Frank-Wolfe Algorithm under ParallelizationBenjamin Dubois-Taine, Francis R. Bach, Quentin Berthet, Adrien B. TaylorNeurIPS 2022 · 被引用 6 次
- A Bregman Proximal Stochastic Gradient Method with Extrapolation for Nonconvex Nonsmooth ProblemsQingsong Wang, Zehui Liu, Chunfeng Cui, Deren HanAAAI 2024 · 被引用 6 次
