Exponential-Family Harmoniums with Neural Sufficient Statistics
Azwar Abdulsalam, Joseph G. Makin
Abstract
Exponential-family harmoniums (EFHs) generalize the restricted Boltzmann machine beyond Bernoulli random variables to other exponential families. Here we show how to extend the EFH beyond standard exponential families (Poisson, Gaussian, etc.), by allowing the sufficient statistics for the hidden units to be arbitrary functions of the observed data, parameterized by deep neural networks. This rules out the standard sampling scheme, block Gibbs sampling, so we replace it with a form of Langevin dynamics within Gibbs, inspired by a recent method for training Gaussian restricted Boltzmann machines (GRBMs). With Gibbs-Langevin, the GRBM can successfully model small datasets like MNIST and CelebA-32, but struggles with CIFAR-10, and cannot scale to larger images because it lacks convolutions. In contrast, our neural-network EFHs (NN-EFHs) generate high-quality samples from CIFAR-10 and scale well to CelebA-HQ. On these datasets, the NN-EFH achieves FID scores that are 25--50% lower than a standard energy-based model with a similar neural-network architecture and the same number of parameters; and competitive with noise-conditional score networks, which utilize more complex neural networks (U-nets) and require considerably more sampling steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based ModelsErik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu et al.AAAI 2020 · 182 citations
- Improved Contrastive Divergence Training of Energy-Based ModelsYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2021 · 171 citations
Related papers
- Expressive probabilistic sampling in recurrent neural networksShirui Chen, Linxing Jiang, Rajesh P. N. Rao, Eric Shea-BrownNeurIPS 2023 · 4 citations
- Bi-level Score Matching for Learning Energy-based Latent Variable ModelsFan Bao, Chongxuan Li, Taufik Xu, Hang Su et al.NeurIPS 2020 · 16 citations
- Score-Based Generative Modeling with Critically-Damped Langevin DiffusionTim Dockhorn, Arash Vahdat, Karsten KreisICLR 2022 · 276 citations
- On Energy-Based Models with Overparametrized Shallow Neural NetworksCarles Domingo-Enrich, Alberto Bietti, Eric Vanden-Eijnden, Joan BrunaICML 2021 · 10 citations
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
