Training Deep Energy-Based Models with f-Divergence Minimization
Lantao Yu, Yang Song, Jiaming Song, Stefano Ermon
Abstract
Deep energy-based models (EBMs) are very flexible in distribution parametrization but computationally challenging because of the intractable partition function. They are typically trained via maximum likelihood, using contrastive divergence to approximate the gradient of the KL divergence between data and model distribution. While KL divergence has many desirable properties, other f -divergences have shown advantages in training implicit density generative models such as generative adversarial networks. In this paper, we propose a general variational framework termed f -EBM to train EBMs using any desired f -divergence. We introduce a corresponding optimization algorithm and prove its local convergence property with non-linear dynamical systems theory. Experimental results demonstrate the superiority of f -EBM over contrastive divergence, as well as the benefits of training EBMs using f -divergences other than KL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 121e4150-0fd2-46e0-b530-b7bd366391e4Cited by top-tier papers22
- Generalized Energy Based ModelsMichael Arbel, Liang Zhou, Arthur GrettonICLR 2021 · 254 citations
- Improved Contrastive Divergence Training of Energy-Based ModelsYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2021 · 171 citations
- Telescoping Density-Ratio EstimationBenjamin Rhodes, Kai Xu, Michael U. GutmannNeurIPS 2020 · 148 citations
- VAEBM: A Symbiosis between Variational Autoencoders and Energy-based ModelsZhisheng Xiao, Karsten Kreis, Jan Kautz, Arash VahdatICLR 2021 · 139 citations
- Generative Flow Networks for Discrete Probabilistic ModelingDinghuai Zhang, Nikolay Malkin, Zhen Liu, Alexandra Volokhova et al.ICML 2022 · 131 citations
Builds on1
Related papers
- Pseudo-Spherical Contrastive DivergenceLantao Yu, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 9 citations
- Guiding Energy-based Models via Contrastive Latent VariablesHankook Lee, Jongheon Jeong, Sejun Park, Jinwoo ShinICLR 2023 · 4 citations
- Unbiased Contrastive Divergence Algorithm for Training Energy-Based Latent Variable ModelsYixuan Qiu, Lingsong Zhang, Xiao WangICLR 2020 · 26 citations
- Improving Adversarial Energy-Based Model via Diffusion ProcessCong Geng, Tian Han, Peng-Tao Jiang, Hao Zhang et al.ICML 2024 · 5 citations
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 21 citations
