Efficient Training of Energy-Based Models Using Jarzynski Equality
Davide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-Eijnden
摘要
Energy-based models (EBMs) are generative models inspired by statistical physics with a wide range of applications in unsupervised learning. Their performance is well measured by the cross-entropy (CE) of the model distribution relative to the data distribution. Using the CE as the objective for training is however challenging because the computation of its gradient with respect to the model parameters requires sampling the model distribution. Here we show how results for nonequilibrium thermodynamics based on Jarzynski equality together with tools from sequential Monte-Carlo sampling can be used to perform this computation efficiently and avoid the uncontrolled approximations made using the standard contrastive divergence algorithm. Specifically, we introduce a modification of the unadjusted Langevin algorithm (ULA) in which each walker acquires a weight that enables the estimation of the gradient of the cross-entropy at any step during GD, thereby bypassing sampling biases induced by slow mixing of ULA. We illustrate these results with numerical experiments on Gaussian mixture distributions as well as the MNIST and CIFAR-10 datasets. We show that the proposed approach outperforms methods based on the contrastive divergence algorithm in all the considered situations. Probabilistic models have become a key tool in generative artificial intelligence (AI) and unsupervised learning. Their goal is twofold: explain the training data, and allow the synthesis of new samples. Many flavors have been introduced in the last decades, including variational auto-encoders [1, 2, 3] generative adversarial networks [4, 5] , normalizing flows [6, 7, 8, 9, 10] , diffusion-based models [11, 12, 13] , restricted Boltzmann machines [14, 15, 16] , and energy-based models (EBMs) [17, 18, 19] . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Complexity Analysis of Normalizing Constant Estimation: from Jarzynski Equality to Annealed Importance Sampling and beyondWei Guo, Molei Tao, Yongxin ChenICLR 2026 · 被引用 12 次
- A Diffusive Classification Loss for Learning Energy-based Generative ModelsRuiKang OuYang, Louis Grenioux, Jose Miguel Hernandez-LobatoICML 2026 · 被引用 5 次
- Learning Latent Variable Models via Jarzynski-adjusted Langevin AlgorithmJames Cuin, Davide Carbone, O. Deniz AkyildizNeurIPS 2025 · 被引用 4 次
- Generalizable Reasoning through Compositional Energy MinimizationAlexandru Oarga, Yilun DuNeurIPS 2025 · 被引用 3 次
- Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood MaximizationMikhail Persiianov, Arip Asadulaev, Nikita Andreev, Nikita Starodubcev 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 被引用 958 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Riemannian Continuous Normalizing FlowsEmile Mathieu, Maximilian NickelNeurIPS 2020 · 被引用 198 次
相关 Paper
- Unbiased Contrastive Divergence Algorithm for Training Energy-Based Latent Variable ModelsYixuan Qiu, Lingsong Zhang, Xiao WangICLR 2020 · 被引用 26 次
- Learning Energy-Based Model with Variational Auto-Encoder as Amortized SamplerJianwen Xie, Zilong Zheng, Ping LiAAAI 2021 · 被引用 57 次
- Training Deep Energy-Based Models with f-Divergence MinimizationLantao Yu, Yang Song, Jiaming Song, Stefano ErmonICML 2020 · 被引用 50 次
- Bi-level Score Matching for Learning Energy-based Latent Variable ModelsFan Bao, Chongxuan Li, Taufik Xu, Hang Su 等NeurIPS 2020 · 被引用 16 次
- Guiding Energy-based Models via Contrastive Latent VariablesHankook Lee, Jongheon Jeong, Sejun Park, Jinwoo ShinICLR 2023 · 被引用 4 次
