Structured Stochastic Gradient MCMC
Antonios Alexos, Alex J. Boyd, Stephan Mandt
摘要
Stochastic gradient Markov Chain Monte Carlo (SGMCMC) is considered the gold standard for Bayesian inference in large-scale models, such as Bayesian neural networks. Since practitioners face speed versus accuracy tradeoffs in these models, variational inference (VI) is often the preferable option. Unfortunately, VI makes strong assumptions on both the factorization and functional form of the posterior. In this work, we propose a new non-parametric variational approximation that makes no assumptions about the approximate posterior's functional form and allows practitioners to specify the exact dependencies the algorithm should respect or break. The approach relies on a new Langevin-type algorithm that operates on a modified energy function, where parts of the latent variables are averaged over samples from earlier iterations of the Markov chain. This way, statistical dependencies can be broken in a controlled way, allowing the chain to mix faster. This scheme can be further modified in a"dropout"manner, leading to even more scalability. We test our scheme for ResNet-20 on CIFAR-10, SVHN, and FMNIST. In all cases, we find improvements in convergence speed and/or final accuracy compared to SG-MCMC and VI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Sparse Inducing Points in Deep Gaussian Processes: Enhancing Modeling with Denoising Diffusion Variational InferenceJian Xu, Delu Zeng, John W. PaisleyICML 2024 · 被引用 16 次
- Scalable Bayesian Learning with posteriorsSamuel Duffield, Kaelan Donatella, Johnathan Chiu, Phoebe Klett 等ICLR 2025 · 被引用 2 次
- Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte CarloYu Wang, Jie Ding, Jonathan HugginsICML 2026 · 被引用 1 次
- Bridging the Gap between Variational Inference and Stochastic Gradient MCMC in Function SpaceMengjing Wu, Junyu Xuan, Jie LuICLR 2025
- Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation ModelsNgoc-Quan Pham, Tuan Truong, Quyen Tran, Tan Minh Nguyen 等ICML 2025
它引用的顶会 Paper5
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski 等ICML 2020 · 被引用 409 次
- Sparse Uncertainty Representation in Deep Learning with Inducing WeightsHippolyt Ritter, Martin Kukla, Cheng Zhang, Yingzhen LiNeurIPS 2021 · 被引用 23 次
- Black-Box Variational Inference as a Parametric Approximation to Langevin DynamicsMatthew D. Hoffman, Yian MaICML 2020 · 被引用 16 次
- Embedded-model flows: Combining the inductive biases of model-free deep learning and explicit probabilistic modelingGianluigi Silvestri, Emily Fertig, Dave Moore, Luca AmbrogioniICLR 2022 · 被引用 4 次
相关 Paper
- Bayesian Posterior Approximation With Stochastic EnsemblesOleksandr Balabanov, Bernhard Mehlig, Hampus LinanderCVPR 2023
- Parameter Expanded Stochastic Gradient Markov Chain Monte CarloHyunsu Kim, Giung Nam, Chulhee Yun, Hongseok Yang 等ICLR 2025
- GFlowOut: Dropout with Generative Flow NetworksDianbo Liu, Moksh Jain, Bonaventure F. P. Dossou, Qianli Shen 等ICML 2023 · 被引用 27 次
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 被引用 5 次
- BayesDAG: Gradient-Based Posterior Inference for Causal DiscoveryYashas Annadani, Nick Pawlowski, Joel Jennings, Stefan Bauer 等NeurIPS 2023 · 被引用 54 次
