Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without Sampling
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Richard S. Zemel
Abstract
We present a new method for evaluating and training unnormalized density models. Our approach only requires access to the gradient of the unnormalized model's log-density. We estimate the Stein discrepancy between the data density and the model density defined by a vector function of the data. We parameterize this function with a neural network and fit its parameters to maximize the discrepancy. This yields a novel goodness-of-fit test which outperforms existing methods on high dimensional data. Furthermore, optimizing to minimize this discrepancy produces a novel method for training unnormalized models which scales more gracefully than existing methods. The ability to both learn and compare models is a unique feature of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers45
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- AdaptDiffuser: Diffusion Models as Adaptive Self-evolving PlannersZhixuan Liang, Yao Mu, Mingyu Ding, Fei Ni et al.ICML 2023 · 165 citations
- VAEBM: A Symbiosis between Variational Autoencoders and Energy-based ModelsZhisheng Xiao, Karsten Kreis, Jan Kautz, Arash VahdatICLR 2021 · 139 citations
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
Builds on5
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based ModelsErik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu et al.AAAI 2020 · 182 citations
- Energy-based models for atomic-resolution protein conformationsYilun Du, Joshua Meier, Jerry Ma, Rob Fergus et al.ICLR 2020 · 60 citations
- The Shape of Data: Intrinsic Distance for Data DistributionsAnton Tsitsulin, Marina Munkhoeva, Davide Mottin, Panagiotis Karras et al.ICLR 2020 · 57 citations
- Flow Contrastive Estimation of Energy-Based ModelsRuiqi Gao, Erik Nijkamp, Diederik P. Kingma, Zhen Xu et al.CVPR 2020
Related papers
- A Kernel Stein Test of Goodness of Fit for Sequential ModelsJerome Baum, Heishiro Kanagawa, Arthur GrettonICML 2023 · 12 citations
- A Kernelised Stein Statistic for Assessing Implicit Generative ModelsWenkai Xu, Gesine D. ReinertNeurIPS 2022 · 4 citations
- Sliced Kernelized Stein DiscrepancyWenbo Gong, Yingzhen Li, José Miguel Hernández-LobatoICLR 2021 · 14 citations
- Concrete Score Matching: Generalized Score Matching for Discrete DataChenlin Meng, Kristy Choi, Jiaming Song, Stefano ErmonNeurIPS 2022 · 168 citations
- Approximate Bayesian Inference with Stein Functional Variational Gradient DescentTobias Pielok, Bernd Bischl, David RügamerICLR 2023
