Oops I Took A Gradient: Scalable Sampling for Discrete Distributions
Will Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud, Chris J. Maddison
Abstract
We propose a general and scalable approximate sampling strategy for probabilistic models with discrete variables. Our approach uses gradients of the likelihood function with respect to its discrete inputs to propose updates in a Metropolis-Hastings sampler. We show empirically that this approach outperforms generic samplers in a number of difficult settings including Ising models, Potts models, restricted Boltzmann machines, and factorial hidden Markov models. We also demonstrate the use of our improved sampler for training deep energy-based models (EBM) on high dimensional discrete data. This approach outperforms variational auto-encoders and existing energy-based models. Finally, we give bounds showing that our approach is near-optimal in the class of samplers which propose local updates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers56
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun et al.NeurIPS 2022 · 316 citations
- Automatically Auditing Large Language Models via Discrete OptimizationErik Jones, Anca D. Dragan, Aditi Raghunathan, Jacob SteinhardtICML 2023 · 232 citations
- Concrete Score Matching: Generalized Score Matching for Discrete DataChenlin Meng, Kristy Choi, Jiaming Song, Stefano ErmonNeurIPS 2022 · 168 citations
- Generative Flow Networks for Discrete Probabilistic ModelingDinghuai Zhang, Nikolay Malkin, Zhen Liu, Alexandra Volokhova et al.ICML 2022 · 131 citations
Builds on6
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- On the Anatomy of MCMC-Based Maximum Likelihood Learning of Energy-Based ModelsErik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu et al.AAAI 2020 · 182 citations
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam et al.ICLR 2020 · 147 citations
- Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without SamplingWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICML 2020 · 93 citations
- Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based ModelsMitch Hill, Jonathan Craig Mitchell, Song-Chun ZhuICLR 2021 · 93 citations
Related papers
- Path Auxiliary Proposal for MCMC in Discrete SpaceHaoran Sun, Hanjun Dai, Wei Xia, Arun RamamurthyICLR 2022 · 27 citations
- A Langevin-like Sampler for Discrete DistributionsRuqi Zhang, Xingchao Liu, Qiang LiuICML 2022 · 51 citations
- Gradient-Guided Importance Sampling for Learning Binary Energy-Based ModelsMeng Liu, Haoran Liu, Shuiwang JiICLR 2023
- Undirected Graphical Models as Approximate PosteriorsArash Vahdat, Evgeny Andriyash, William G. MacreadyICML 2020 · 15 citations
- No MCMC for me: Amortized sampling for fast and stable training of energy-based modelsWill Sussman Grathwohl, Jacob Jin Kelly, Milad Hashemi, Mohammad Norouzi et al.ICLR 2021 · 75 citations
