Contrastive Divergence Learning is a Time Reversal Adversarial Game
Omer Yair, Tomer Michaeli
Abstract
Contrastive divergence (CD) learning is a classical method for fitting unnormalized statistical models to data samples. Despite its wide-spread use, the convergence properties of this algorithm are still not well understood. The main source of difficulty is an unjustified approximation which has been used to derive the gradient of the loss. In this paper, we present an alternative derivation of CD that does not require any approximation and sheds new light on the objective that is actually being optimized by the algorithm. Specifically, we show that CD is an adversarial learning procedure, where a discriminator attempts to classify whether a Markov chain generated from the model has been time-reversed. Thus, although predating generative adversarial networks (GANs) by more than a decade, CD is, in fact, closely related to these techniques. Our derivation settles well with previous observations, which have concluded that CD's update steps cannot be expressed as the gradients of any fixed objective function. In addition, as a byproduct, our derivation reveals a simple correction that can be used as an alternative to Metropolis-Hastings rejection, which is required when the underlying Markov chain is inexact (e.g., when using Langevin dynamics with a large step).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44e13a1f-403e-438a-b8bc-bfa93c29d806Cited by top-tier papers3
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 21 citations
- Energy Discrepancies: A Score-Independent Loss for Energy-Based ModelsTobias Schröder, Zijing Ou, Jen Lim, Yingzhen Li et al.NeurIPS 2023 · 15 citations
- Are GANs overkill for NLP?David Alvarez-Melis, Vikas Garg, Adam KalaiNeurIPS 2022 · 2 citations
Builds on2
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Unbiased Contrastive Divergence Algorithm for Training Energy-Based Latent Variable ModelsYixuan Qiu, Lingsong Zhang, Xiao WangICLR 2020 · 26 citations
Related papers
- MonoFlow: Rethinking Divergence GANs via the Perspective of Wasserstein Gradient FlowsMingxuan Yi, Zhanxing Zhu, Song LiuICML 2023 · 18 citations
- Training GANs with Stronger Augmentations via Contrastive DiscriminatorJongheon Jeong, Jinwoo ShinICLR 2021 · 68 citations
- Deep MMD Gradient Flow without adversarial trainingAlexandre Galashov, Valentin De Bortoli, Arthur GrettonICLR 2025 · 1 citation
- Near-Optimality of Contrastive Divergence AlgorithmsPierre Glaser, Kevin Han Huang, Arthur GrettonNeurIPS 2024 · 2 citations
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 5 citations
