On the Convergence of Hamiltonian Monte Carlo with Stochastic Gradients
Difan Zou, Quanquan Gu
Abstract
Hamiltonian Monte Carlo (HMC), built based on the Hamilton's equation, has been witnessed great success in sampling from high-dimensional posterior distributions. However, it also suffers from computational inefficiency, especially for large training datasets. One common idea to overcome this computational bottleneck is using stochastic gradients, which only queries a mini-batch of training data in each iteration. However, unlike the extensive studies on the convergence analysis of HMC using full gradients, few works focus on establishing the convergence guarantees of stochastic gradient HMC algorithms. In this paper, we propose a general framework for proving the convergence rate of HMC with stochastic gradient estimators, for sampling from strongly log-concave and log-smooth target distributions. We show that the convergence to the target distribution in 2-Wasserstein distance can be guaranteed as long as the stochastic gradient estimator is unbiased and its variance is upper bounded along the algorithm trajectory. We further apply the proposed framework to analyze the convergence rates of HMC with four standard stochastic gradient estimators: mini-batch stochastic gradient (SG), stochastic variance reduced gradient (SVRG), stochastic average gradient (SAGA), and control variate gradient (CVG). Theoretical results explain the inefficiency of mini-batch SG, and suggest that SVRG and SAGA perform better in the tasks with high-precision requirements, while CVG performs better for large dataset. Experiment results verify our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1dcd18d-ec15-46f7-80ea-52b303051afcCited by top-tier papers5
- A Symmetry-Aware Exploration of Bayesian Neural Network PosteriorsOlivier Laurent, Emanuel Aldea, Gianni FranchiICLR 2024 · 12 citations
- Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin DynamicsHaoyang Zheng, Hengrong Du, Qi Feng, Wei Deng et al.ICML 2024 · 9 citations
- Faster Sampling via Stochastic Gradient Proximal SamplerXunpeng Huang, Difan Zou, Hanze Dong, Yian Ma et al.ICML 2024 · 4 citations
- Revisiting the Effects of Stochasticity for Hamiltonian SamplersGiulio Franzese, Dimitrios Milios, Maurizio Filippone, Pietro MichiardiICML 2022 · 3 citations
- Accelerating Hamiltonian Monte Carlo via Chebyshev Integration TimeJun-Kun Wang, Andre WibisonoICLR 2023
Builds on2
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 58 citations
- Non-convex Learning via Replica Exchange Stochastic Gradient MCMCWei Deng, Qi Feng, Liyao Gao, Faming Liang et al.ICML 2020 · 54 citations
Related papers
- A Hybrid Stochastic Gradient Hamiltonian Monte Carlo MethodChao Zhang, Zhijian Li, Zebang Shen, Jiahao Xie et al.AAAI 2021 · 3 citations
- Variance Reduction in Stochastic Particle-Optimization SamplingJianyi Zhang, Yang Zhao, Changyou ChenICML 2020 · 13 citations
- Stochastic Reweighted Gradient DescentAyoub El Hanchi, David A. Stephens, Chris J. MaddisonICML 2022 · 10 citations
- Accelerating Langevin Monte Carlo via Efficient Stochastic Runge-Kutta Methods beyond Log-ConcavityBin Yang, Xiaojie WangICML 2026 · 1 citation
- A Gradient Based Strategy for Hamiltonian Monte Carlo Hyperparameter OptimizationAndrew Campbell, Wenlong Chen, Vincent Stimper, José Miguel Hernández-Lobato et al.ICML 2021 · 20 citations
