Learning Energy-based Model via Dual-MCMC Teaching
Jiali Cui, Tian Han
Abstract
This paper studies the fundamental learning problem of the energy-based model (EBM). Learning the EBM can be achieved using the maximum likelihood estimation (MLE), which typically involves the Markov Chain Monte Carlo (MCMC) sampling, such as the Langevin dynamics. However, the noise-initialized Langevin dynamics can be challenging in practice and hard to mix. This motivates the exploration of joint training with the generator model where the generator model serves as a complementary model to bypass MCMC sampling. However, such a method can be less accurate than the MCMC and result in biased EBM learning. While the generator can also serve as an initializer model for better MCMC sampling, its learning can be biased since it only matches the EBM and has no access to empirical training examples. Such biased generator learning may limit the potential of learning the EBM. To address this issue, we present a joint learning framework that interweaves the maximum likelihood learning algorithm for both the EBM and the complementary generator model. In particular, the generator model is learned by MLE to match both the EBM and the empirical data distribution, making it a more informative initializer for MCMC sampling of EBM. Learning generator with observed examples typically requires inference of the generator posterior. To ensure accurate and efficient inference, we adopt the MCMC posterior sampling and introduce a complementary inference model to initialize such latent MCMC sampling. We show that three separate models can be seamlessly integrated into our joint framework through two (dual-) MCMC teaching, enabling effective and efficient EBM learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 941a9139-3832-455e-a4fb-b06cd844794fCited by top-tier papers4
- Learning Latent Space Hierarchical EBM Diffusion ModelsJiali Cui, Tian HanICML 2024 · 7 citations
- Improving Adversarial Energy-Based Model via Diffusion ProcessCong Geng, Tian Han, Peng-Tao Jiang, Hao Zhang et al.ICML 2024 · 5 citations
- Diffusion Federated DatasetSeok-Ju Hahn, Junghye LeeNeurIPS 2025 · 3 citations
- Latent-Guided Cooperative Energy-Based ModelsCong Geng, Xue Han, Ye Yuan, Qiang Hu et al.ICML 2026
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
Related papers
- Learning Energy-Based Model with Variational Auto-Encoder as Amortized SamplerJianwen Xie, Zilong Zheng, Ping LiAAAI 2021 · 57 citations
- Learning Joint Latent Space EBM Prior Model for Multi-layer GeneratorJiali Cui, Ying Nian Wu, Tian HanCVPR 2023
- Learning Hierarchical Features with Joint Latent Space Energy-Based PriorJiali Cui, Ying Nian Wu, Tian HanICCV 2023 · 11 citations
- A Tale of Two Flows: Cooperative Learning of Langevin Flow and Normalizing Flow Toward Energy-Based ModelJianwen Xie, Yaxuan Zhu, Jun Li, Ping LiICLR 2022 · 53 citations
- No MCMC for me: Amortized sampling for fast and stable training of energy-based modelsWill Sussman Grathwohl, Jacob Jin Kelly, Milad Hashemi, Mohammad Norouzi et al.ICLR 2021 · 75 citations
