Differentiable Annealed Importance Sampling and the Perils of Gradient Noise
Guodong Zhang, Kyle Hsu, Jianing Li, Chelsea Finn, Roger B. Grosse
Abstract
Annealed importance sampling (AIS) and related algorithms are highly effective tools for marginal likelihood estimation, but are not fully differentiable due to the use of Metropolis-Hastings correction steps. Differentiability is a desirable property as it would admit the possibility of optimizing marginal likelihood as an objective using gradient-based methods. To this end, we propose Differentiable AIS (DAIS), a variant of AIS which ensures differentiability by abandoning the Metropolis-Hastings corrections. As a further advantage, DAIS allows for mini-batch gradients. We provide a detailed convergence analysis for Bayesian linear regression which goes beyond previous analyses by explicitly accounting for the sampler not having reached equilibrium. Using this analysis, we prove that DAIS is consistent in the full-batch setting and provide a sublinear convergence rate. Furthermore, motivated by the problem of learning from large-scale datasets, we study a stochastic variant of DAIS that uses mini-batch gradients. Surprisingly, stochastic DAIS can be arbitrarily bad due to a fundamental incompatibility between the goals of last-iterate convergence to the posterior and elimination of the accumulated stochastic error. This is in stark contrast with other settings such as gradient-based optimization and Langevin dynamics, where the effect of gradient noise can be washed out by taking smaller steps. This indicates that annealing-based marginal likelihood estimation with stochastic gradients may require new ideas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94c27d1d-8f4f-444f-9f4d-87f684c3fb81Cited by top-tier papers20
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum et al.ICML 2022 · 83 citations
- Continual Repeated Annealed Flow Transport Monte CarloAlexander G. de G. Matthews, Michael Arbel, Danilo Jimenez Rezende, Arnaud DoucetICML 2022 · 69 citations
- Score-Based Diffusion meets Annealed Importance SamplingArnaud Doucet, Will Grathwohl, Alexander G. de G. Matthews, Heiko StrathmannNeurIPS 2022 · 68 citations
- Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for SamplingDenis Blessing, Xiaogang Jia, Johannes Esslinger, Francisco Vargas et al.ICML 2024 · 47 citations
- MCMC Variational Inference via Uncorrected Hamiltonian AnnealingTomas Geffner, Justin DomkeNeurIPS 2021 · 45 citations
Builds on5
- Monte Carlo Variational Auto-EncodersAchille Thin, Nikita Kotelevskii, Arnaud Doucet, Alain Durmus et al.ICML 2021 · 51 citations
- MCMC Variational Inference via Uncorrected Hamiltonian AnnealingTomas Geffner, Justin DomkeNeurIPS 2021 · 45 citations
- Evaluating Lossy Compression Rates of Deep Generative ModelsSicong Huang, Alireza Makhzani, Yanshuai Cao, Roger B. GrosseICML 2020 · 30 citations
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable ModelsYucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu et al.ICLR 2020 · 29 citations
- Improving Lossless Compression Rates via Monte Carlo Bits-Back CodingYangjun Ruan, Karen Ullrich, Daniel Severo, James Townsend et al.ICML 2021 · 25 citations
Related papers
- Differentiable Annealed Importance Sampling Minimizes The Jensen-Shannon Divergence Between Initial and Target DistributionJohannes Zenn, Robert BamlerICML 2024
- Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?Emanuel Sommer, Kangning Diao, Jakob Robnik, Uros Seljak et al.ICML 2026 · 6 citations
- Revisiting the Effects of Stochasticity for Hamiltonian SamplersGiulio Franzese, Dimitrios Milios, Maurizio Filippone, Pietro MichiardiICML 2022 · 3 citations
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 5 citations
- Practical and Scalable Hamiltonian Monte Carlo Without the Metropolis TestJakob Robnik, Reuben Cohn-Gordon, Uros SeljakICML 2026 · 5 citations
