Differentiable Annealed Importance Sampling and the Perils of Gradient Noise
Guodong Zhang, Kyle Hsu, Jianing Li, Chelsea Finn, Roger B. Grosse
摘要
Annealed importance sampling (AIS) and related algorithms are highly effective tools for marginal likelihood estimation, but are not fully differentiable due to the use of Metropolis-Hastings correction steps. Differentiability is a desirable property as it would admit the possibility of optimizing marginal likelihood as an objective using gradient-based methods. To this end, we propose Differentiable AIS (DAIS), a variant of AIS which ensures differentiability by abandoning the Metropolis-Hastings corrections. As a further advantage, DAIS allows for mini-batch gradients. We provide a detailed convergence analysis for Bayesian linear regression which goes beyond previous analyses by explicitly accounting for the sampler not having reached equilibrium. Using this analysis, we prove that DAIS is consistent in the full-batch setting and provide a sublinear convergence rate. Furthermore, motivated by the problem of learning from large-scale datasets, we study a stochastic variant of DAIS that uses mini-batch gradients. Surprisingly, stochastic DAIS can be arbitrarily bad due to a fundamental incompatibility between the goals of last-iterate convergence to the posterior and elimination of the accumulated stochastic error. This is in stark contrast with other settings such as gradient-based optimization and Langevin dynamics, where the effect of gradient noise can be washed out by taking smaller steps. This indicates that annealing-based marginal likelihood estimation with stochastic gradients may require new ideas.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum 等ICML 2022 · 被引用 83 次
- Continual Repeated Annealed Flow Transport Monte CarloAlexander G. de G. Matthews, Michael Arbel, Danilo Jimenez Rezende, Arnaud DoucetICML 2022 · 被引用 69 次
- Score-Based Diffusion meets Annealed Importance SamplingArnaud Doucet, Will Grathwohl, Alexander G. de G. Matthews, Heiko StrathmannNeurIPS 2022 · 被引用 68 次
- Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for SamplingDenis Blessing, Xiaogang Jia, Johannes Esslinger, Francisco Vargas 等ICML 2024 · 被引用 47 次
- MCMC Variational Inference via Uncorrected Hamiltonian AnnealingTomas Geffner, Justin DomkeNeurIPS 2021 · 被引用 45 次
它引用的顶会 Paper5
- Monte Carlo Variational Auto-EncodersAchille Thin, Nikita Kotelevskii, Arnaud Doucet, Alain Durmus 等ICML 2021 · 被引用 51 次
- MCMC Variational Inference via Uncorrected Hamiltonian AnnealingTomas Geffner, Justin DomkeNeurIPS 2021 · 被引用 45 次
- Evaluating Lossy Compression Rates of Deep Generative ModelsSicong Huang, Alireza Makhzani, Yanshuai Cao, Roger B. GrosseICML 2020 · 被引用 30 次
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable ModelsYucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu 等ICLR 2020 · 被引用 29 次
- Improving Lossless Compression Rates via Monte Carlo Bits-Back CodingYangjun Ruan, Karen Ullrich, Daniel Severo, James Townsend 等ICML 2021 · 被引用 25 次
相关 Paper
- Differentiable Annealed Importance Sampling Minimizes The Jensen-Shannon Divergence Between Initial and Target DistributionJohannes Zenn, Robert BamlerICML 2024
- Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?Emanuel Sommer, Kangning Diao, Jakob Robnik, Uros Seljak 等ICML 2026 · 被引用 6 次
- Revisiting the Effects of Stochasticity for Hamiltonian SamplersGiulio Franzese, Dimitrios Milios, Maurizio Filippone, Pietro MichiardiICML 2022 · 被引用 3 次
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 被引用 5 次
- Practical and Scalable Hamiltonian Monte Carlo Without the Metropolis TestJakob Robnik, Reuben Cohn-Gordon, Uros SeljakICML 2026 · 被引用 5 次
