Training Deep Energy-Based Models with f-Divergence Minimization
Lantao Yu, Yang Song, Jiaming Song, Stefano Ermon
摘要
Deep energy-based models (EBMs) are very flexible in distribution parametrization but computationally challenging because of the intractable partition function. They are typically trained via maximum likelihood, using contrastive divergence to approximate the gradient of the KL divergence between data and model distribution. While KL divergence has many desirable properties, other f -divergences have shown advantages in training implicit density generative models such as generative adversarial networks. In this paper, we propose a general variational framework termed f -EBM to train EBMs using any desired f -divergence. We introduce a corresponding optimization algorithm and prove its local convergence property with non-linear dynamical systems theory. Experimental results demonstrate the superiority of f -EBM over contrastive divergence, as well as the benefits of training EBMs using f -divergences other than KL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Generalized Energy Based ModelsMichael Arbel, Liang Zhou, Arthur GrettonICLR 2021 · 被引用 254 次
- Improved Contrastive Divergence Training of Energy-Based ModelsYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2021 · 被引用 171 次
- Telescoping Density-Ratio EstimationBenjamin Rhodes, Kai Xu, Michael U. GutmannNeurIPS 2020 · 被引用 148 次
- VAEBM: A Symbiosis between Variational Autoencoders and Energy-based ModelsZhisheng Xiao, Karsten Kreis, Jan Kautz, Arash VahdatICLR 2021 · 被引用 139 次
- Generative Flow Networks for Discrete Probabilistic ModelingDinghuai Zhang, Nikolay Malkin, Zhen Liu, Alexandra Volokhova 等ICML 2022 · 被引用 131 次
它引用的顶会 Paper1
相关 Paper
- Pseudo-Spherical Contrastive DivergenceLantao Yu, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 被引用 9 次
- Guiding Energy-based Models via Contrastive Latent VariablesHankook Lee, Jongheon Jeong, Sejun Park, Jinwoo ShinICLR 2023 · 被引用 4 次
- Unbiased Contrastive Divergence Algorithm for Training Energy-Based Latent Variable ModelsYixuan Qiu, Lingsong Zhang, Xiao WangICLR 2020 · 被引用 26 次
- Improving Adversarial Energy-Based Model via Diffusion ProcessCong Geng, Tian Han, Peng-Tao Jiang, Hao Zhang 等ICML 2024 · 被引用 5 次
- Efficient Training of Energy-Based Models Using Jarzynski EqualityDavide Carbone, Mengjian Hua, Simon Coste, Eric Vanden-EijndenNeurIPS 2023 · 被引用 21 次
