Structured Stochastic Gradient MCMC
Antonios Alexos, Alex J. Boyd, Stephan Mandt
Abstract
Stochastic gradient Markov Chain Monte Carlo (SGMCMC) is considered the gold standard for Bayesian inference in large-scale models, such as Bayesian neural networks. Since practitioners face speed versus accuracy tradeoffs in these models, variational inference (VI) is often the preferable option. Unfortunately, VI makes strong assumptions on both the factorization and functional form of the posterior. In this work, we propose a new non-parametric variational approximation that makes no assumptions about the approximate posterior's functional form and allows practitioners to specify the exact dependencies the algorithm should respect or break. The approach relies on a new Langevin-type algorithm that operates on a modified energy function, where parts of the latent variables are averaged over samples from earlier iterations of the Markov chain. This way, statistical dependencies can be broken in a controlled way, allowing the chain to mix faster. This scheme can be further modified in a"dropout"manner, leading to even more scalability. We test our scheme for ResNet-20 on CIFAR-10, SVHN, and FMNIST. In all cases, we find improvements in convergence speed and/or final accuracy compared to SG-MCMC and VI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e171828e-74ae-46a2-b6b8-62710bd69179Cited by top-tier papers6
- Sparse Inducing Points in Deep Gaussian Processes: Enhancing Modeling with Denoising Diffusion Variational InferenceJian Xu, Delu Zeng, John W. PaisleyICML 2024 · 16 citations
- Scalable Bayesian Learning with posteriorsSamuel Duffield, Kaelan Donatella, Johnathan Chiu, Phoebe Klett et al.ICLR 2025 · 2 citations
- Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte CarloYu Wang, Jie Ding, Jonathan HugginsICML 2026 · 1 citation
- Bridging the Gap between Variational Inference and Stochastic Gradient MCMC in Function SpaceMengjing Wu, Junyu Xuan, Jie LuICLR 2025
- Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation ModelsNgoc-Quan Pham, Tuan Truong, Quyen Tran, Tan Minh Nguyen et al.ICML 2025
Builds on5
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Sparse Uncertainty Representation in Deep Learning with Inducing WeightsHippolyt Ritter, Martin Kukla, Cheng Zhang, Yingzhen LiNeurIPS 2021 · 23 citations
- Black-Box Variational Inference as a Parametric Approximation to Langevin DynamicsMatthew D. Hoffman, Yian MaICML 2020 · 16 citations
- Embedded-model flows: Combining the inductive biases of model-free deep learning and explicit probabilistic modelingGianluigi Silvestri, Emily Fertig, Dave Moore, Luca AmbrogioniICLR 2022 · 4 citations
Related papers
- Bayesian Posterior Approximation With Stochastic EnsemblesOleksandr Balabanov, Bernhard Mehlig, Hampus LinanderCVPR 2023
- Parameter Expanded Stochastic Gradient Markov Chain Monte CarloHyunsu Kim, Giung Nam, Chulhee Yun, Hongseok Yang et al.ICLR 2025
- GFlowOut: Dropout with Generative Flow NetworksDianbo Liu, Moksh Jain, Bonaventure F. P. Dossou, Qianli Shen et al.ICML 2023 · 27 citations
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 5 citations
- BayesDAG: Gradient-Based Posterior Inference for Causal DiscoveryYashas Annadani, Nick Pawlowski, Joel Jennings, Stefan Bauer et al.NeurIPS 2023 · 54 citations
