On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based Models
Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, Ying Nian Wu
Abstract
This study investigates the effects of Markov chain Monte Carlo (MCMC) sampling in unsupervised Maximum Likelihood (ML) learning. Our attention is restricted to the family of unnormalized probability densities for which the negative log density (or energy function) is a ConvNet. We find that many of the techniques used to stabilize training in previous studies are not necessary. ML learning with a ConvNet potential requires only a few hyper-parameters and no regularization. Using this minimal framework, we identify a variety of ML learning outcomes that depend solely on the implementation of MCMC sampling. On one hand, we show that it is easy to train an energy-based model which can sample realistic images with short-run Langevin. ML can be effective and stable even when MCMC samples have much higher energy than true steady-state samples throughout training. Based on this insight, we introduce an ML method with purely noise-initialized MCMC, high-quality short-run synthesis, and the same budget as ML with informative MCMC initialization such as CD or PCD. Unlike previous models, our energy model can obtain realistic high-diversity samples from a noise signal after training. On the other hand, ConvNet potentials learned with non-convergent MCMC do not have a valid steady-state and cannot be considered approximate unnormalized densities of the training data because long-run MCMC samples differ greatly from observed images. We show that it is much harder to train a ConvNet potential to learn a steady-state over realistic images. To our knowledge, long-run MCMC samples of all previous models lose the realism of short-run samples. With correct tuning of Langevin noise, we train the first ConvNet potentials for which long-run and steady-state MCMC samples are realistic images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f213ce8-744d-449b-97e0-8f0a916002f8Cited by top-tier papers67
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- On Aliased Resizing and Surprising Subtleties in GAN EvaluationGaurav Parmar, Richard Zhang, Jun-Yan ZhuCVPR 2022 · 250 citations
- Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCYilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum et al.ICML 2023 · 219 citations
- COLD Decoding: Energy-based Constrained Text Generation with Langevin DynamicsLianhui Qin, Sean Welleck, Daniel Khashabi, Yejin ChoiNeurIPS 2022 · 217 citations
Related papers
- Learning Energy-based Model via Dual-MCMC TeachingJiali Cui, Tian HanNeurIPS 2023 · 14 citations
- Hamiltonian Dynamics with Non-Newtonian Momentum for Rapid SamplingGreg Ver Steeg, Aram GalstyanNeurIPS 2021 · 18 citations
- A Unified Contrastive Energy-based Model for Understanding the Generative Ability of Adversarial TrainingYifei Wang, Yisen Wang, Jiansheng Yang, Zhouchen LinICLR 2022 · 19 citations
- Explaining the effects of non-convergent MCMC in the training of Energy-Based ModelsElisabeth Agoritsas, Giovanni Catania, Aurélien Decelle, Beatriz SeoaneICML 2023 · 17 citations
- Langevin Autoencoders for Learning Deep Latent Variable ModelsShohei Taniguchi, Yusuke Iwasawa, Wataru Kumagai, Yutaka MatsuoNeurIPS 2022 · 2 citations
