Non-reversible Parallel Tempering for Deep Posterior Approximation
Wei Deng, Qian Zhang, Qi Feng, Faming Liang, Guang Lin
Abstract
Parallel tempering (PT), also known as replica exchange, is the go-to workhorse for simulations of multi-modal distributions. The key to the success of PT is to adopt efficient swap schemes. The popular deterministic even-odd (DEO) scheme exploits the non-reversibility property and has successfully reduced the communication cost from O(P 2 ) to O(P ) given sufficiently many P chains. However, such an innovation largely disappears in big data due to the limited chains and few bias-corrected swaps. To handle this issue, we generalize the DEO scheme to promote non-reversibility and propose a few solutions to tackle the underlying bias caused by the geometric stopping time. Notably, in big data scenarios, we obtain an appealing communication cost O(P log P ) based on the optimal window size. In addition, we also adopt stochastic gradient descent (SGD) with large and constant learning rates as exploration kernels. Such a user-friendly nature enables us to conduct approximation tasks for complex posteriors without much tuning costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52f4ca18-5e59-408c-85c1-d649f93a0ce0Cited by top-tier papers2
- Diffusive Gibbs SamplingWenlin Chen, Mingtian Zhang, Brooks Paige, José Miguel Hernández-Lobato et al.ICML 2024 · 21 citations
- Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin DynamicsHaoyang Zheng, Hengrong Du, Qi Feng, Wei Deng et al.ICML 2024 · 9 citations
Builds on7
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- Non-convex Learning via Replica Exchange Stochastic Gradient MCMCWei Deng, Qi Feng, Liyao Gao, Faming Liang et al.ICML 2020 · 54 citations
- Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear ModelRaphaël Berthier, Francis R. Bach, Pierre GaillardNeurIPS 2020 · 49 citations
- Interacting Contour Stochastic Gradient Langevin DynamicsWei Deng, Siqi Liang, Botao Hao, Guang Lin et al.ICLR 2022 · 13 citations
Related papers
- Accelerated Parallel Tempering via Neural TransportsLeo Zhang, Peter Potaptchik, Jiajun He, Yuanqi Du et al.ICLR 2026 · 14 citations
- Parallel tempering on optimized pathsSaifuddin Syed, Vittorio Romaniello, Trevor Campbell, Alexandre Bouchard-CôtéICML 2021 · 28 citations
- Test-Time Guidance for Flow-Based Generative Models via Parallel Tempering on Source DistributionsShih-Hsin Wang, Joel Keller, Taos Transue, Drake Brown et al.ICML 2026
- Parallel Tempering With a Variational ReferenceNikola Surjanovic, Saifuddin Syed, Alexandre Bouchard-Côté, Trevor CampbellNeurIPS 2022 · 23 citations
- Continuously Tempered PDMP samplersMatthew Sutton, Robert Salomone, Augustin Chevallier, Paul FearnheadNeurIPS 2022 · 2 citations
