Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to Optimization
Yuri Kinoshita, Taiji Suzuki
摘要
The stochastic gradient Langevin Dynamics is one of the most fundamental algorithms to solve sampling problems and non-convex optimization appearing in several machine learning applications. Especially, its variance reduced versions have nowadays gained particular attention. In this paper, we study two variants of this kind, namely, the Stochastic Variance Reduced Gradient Langevin Dynamics and the Stochastic Recursive Gradient Langevin Dynamics. We prove their convergence to the objective distribution in terms of KL-divergence under the sole assumptions of smoothness and Log-Sobolev inequality which are weaker conditions than those used in prior works for these algorithms. With the batch size and the inner loop length set to √ n, the gradient complexity to achieve an -precision is Õ((n + dn 1/2 -1 )γ 2 L 2 α -2 ), which is an improvement from any previous analyses. We also show some essential applications of our result to non-convex optimization. Zou et al. (2019b) Smooth, Dissipative 2-Wass. Õ (n+ n 1/2 2 µ 3/2 * )∧ µ -2 * 4
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Algorithmic Stability of Heavy-Tailed SGD with General Loss FunctionsAnant Raj, Lingjiong Zhu, Mert Gürbüzbalaban, Umut SimsekliICML 2023 · 被引用 21 次
- Time-Independent Information-Theoretic Generalization Bounds for SGLDFutoshi Futami, Masahiro FujisawaNeurIPS 2023 · 被引用 12 次
- Provably Fast Finite Particle Variants of SVGD via Virtual Particle Stochastic ApproximationAniket Das, Dheeraj NagarajNeurIPS 2023 · 被引用 10 次
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 被引用 10 次
- On Divergence Measures for Training GFlowNetsTiago da Silva, Eliezer de Souza da Silva, Diego MesquitaNeurIPS 2024 · 被引用 8 次
相关 Paper
- Sharp Analysis of Stochastic Optimization under Global Kurdyka-Lojasiewicz InequalityIlyas Fatkhullin, Jalal Etesami, Niao He, Negar KiyavashNeurIPS 2022 · 被引用 34 次
- Generalization of noisy SGD in unbounded non-convex settingsLeello Tadesse Dadi, Volkan CevherICML 2025
- Time-independent Generalization Bounds for SGLD in Non-convex SettingsTyler Farghly, Patrick RebeschiniNeurIPS 2021 · 被引用 30 次
- Faster Sampling via Stochastic Gradient Proximal SamplerXunpeng Huang, Difan Zou, Hanze Dong, Yian Ma 等ICML 2024 · 被引用 4 次
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 被引用 95 次
