Contrastive Divergence Learning is a Time Reversal Adversarial Game
Omer Yair, Tomer Michaeli
摘要
Contrastive divergence (CD) learning is a classical method for fitting unnormalized statistical models to data samples. Despite its wide-spread use, the convergence properties of this algorithm are still not well understood. The main source of difficulty is an unjustified approximation which has been used to derive the gradient of the loss. In this paper, we present an alternative derivation of CD that does not require any approximation and sheds new light on the objective that is actually being optimized by the algorithm. Specifically, we show that CD is an adversarial learning procedure, where a discriminator attempts to classify whether a Markov chain generated from the model has been time-reversed. Thus, although predating generative adversarial networks (GANs) by more than a decade, CD is, in fact, closely related to these techniques. Our derivation settles well with previous observations, which have concluded that CD's update steps cannot be expressed as the gradients of any fixed objective function. In addition, as a byproduct, our derivation reveals a simple correction that can be used as an alternative to Metropolis-Hastings rejection, which is required when the underlying Markov chain is inexact (e.g., when using Langevin dynamics with a large step).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 被引用 21 次
- Energy Discrepancies: A Score-Independent Loss for Energy-Based ModelsTobias Schröder, Zijing Ou, Jen Lim, Yingzhen Li 等NeurIPS 2023 · 被引用 15 次
- Are GANs overkill for NLP?David Alvarez-Melis, Vikas Garg, Adam KalaiNeurIPS 2022 · 被引用 2 次
它引用的顶会 Paper2
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Unbiased Contrastive Divergence Algorithm for Training Energy-Based Latent Variable ModelsYixuan Qiu, Lingsong Zhang, Xiao WangICLR 2020 · 被引用 26 次
相关 Paper
- MonoFlow: Rethinking Divergence GANs via the Perspective of Wasserstein Gradient FlowsMingxuan Yi, Zhanxing Zhu, Song LiuICML 2023 · 被引用 18 次
- Training GANs with Stronger Augmentations via Contrastive DiscriminatorJongheon Jeong, Jinwoo ShinICLR 2021 · 被引用 68 次
- Deep MMD Gradient Flow without adversarial trainingAlexandre Galashov, Valentin De Bortoli, Arthur GrettonICLR 2025 · 被引用 1 次
- Near-Optimality of Contrastive Divergence AlgorithmsPierre Glaser, Kevin Han Huang, Arthur GrettonNeurIPS 2024 · 被引用 2 次
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 被引用 5 次
