Explaining the effects of non-convergent MCMC in the training of Energy-Based Models
Elisabeth Agoritsas, Giovanni Catania, Aurélien Decelle, Beatriz Seoane
摘要
In this paper, we quantify the impact of using non-convergent Markov chains to train energy-based models (EBMs). In particular, we show analytically that EBMs trained with non-persistent short runs to estimate the gradient can perfectly reproduce a set of empirical statistics of the data, not at the level of the equilibrium measure, but through a precise dynamical process. Our results provide a first-principles explanation for the observations of recent works proposing the strategy of using short runs starting from random initial conditions as an efficient way to generate high-quality samples in EBMs, and lay the groundwork for using EBMs as diffusion models. After explaining this effect in generic EBMs, we analyze two solvable models in which the effect of the non-convergent sampling in the trained parameters can be described in detail. Finally, we test these predictions numerically on a ConvNet EBM and a Boltzmann machine.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Restoring balance: principled under/oversampling of data for optimal classificationEmanuele Loffredo, Mauro Pastore, Simona Cocco, Rémi MonassonICML 2024 · 被引用 13 次
- A Theoretical Framework For Overfitting In Energy-based ModelingGiovanni Catania, Aurélien Decelle, Cyril Furtlehner, Beatriz SeoaneICML 2025
它引用的顶会 Paper7
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Compositional Visual Generation with Energy Based ModelsYilun Du, Shuang Li, Igor MordatchNeurIPS 2020 · 被引用 225 次
- On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based ModelsErik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu 等AAAI 2020 · 被引用 182 次
- A Tale of Two Flows: Cooperative Learning of Langevin Flow and Normalizing Flow Toward Energy-Based ModelJianwen Xie, Yaxuan Zhu, Jun Li, Ping LiICLR 2022 · 被引用 53 次
相关 Paper
- Equilibrium and non-Equilibrium regimes in the learning of Restricted Boltzmann MachinesAurélien Decelle, Cyril Furtlehner, Beatriz SeoaneNeurIPS 2021 · 被引用 40 次
- Unbiased Contrastive Divergence Algorithm for Training Energy-Based Latent Variable ModelsYixuan Qiu, Lingsong Zhang, Xiao WangICLR 2020 · 被引用 26 次
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 被引用 21 次
- Energy-Based Modelling for Discrete and Mixed Data via Heat Equations on Structured SpacesTobias Schröder, Zijing Ou, Yingzhen Li, Andrew B. DuncanNeurIPS 2024 · 被引用 5 次
- Learning Energy-Based Models by Diffusion Recovery LikelihoodRuiqi Gao, Yang Song, Ben Poole, Ying Nian Wu 等ICLR 2021 · 被引用 144 次
