Gradient-Guided Importance Sampling for Learning Binary Energy-Based Models
Meng Liu, Haoran Liu, Shuiwang Ji
Abstract
Learning energy-based models (EBMs) is known to be difficult especially on discrete data where gradient-based learning strategies cannot be applied directly. Although ratio matching is a sound method to learn discrete EBMs, it suffers from expensive computation and excessive memory requirements, thereby resulting in difficulties in learning EBMs on high-dimensional data. Motivated by these limitations, in this study, we propose ratio matching with gradient-guided importance sampling (RMwGGIS). Particularly, we use the gradient of the energy function w.r.t. the discrete data space to approximately construct the provably optimal proposal distribution, which is subsequently used by importance sampling to efficiently estimate the original ratio matching objective. We perform experiments on density modeling over synthetic discrete data, graph generation, and training Ising models to evaluate our proposed method. The experimental results demonstrate that our method can significantly alleviate the limitations of ratio matching, perform more effectively in practice, and scale to high-dimensional problems. Our implementation is available at https://github.com/divelab/RMwGGIS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72105dce-aa57-4aec-8e71-87ef4ef28e43Cited by top-tier papers2
- Learning Unnormalized Statistical Models via Compositional OptimizationWei Jiang, Jiayu Qin, Lingyu Wu, Changyou Chen et al.ICML 2023 · 8 citations
- Energy-Based Modelling for Discrete and Mixed Data via Heat Equations on Structured SpacesTobias Schröder, Zijing Ou, Yingzhen Li, Andrew B. DuncanNeurIPS 2024 · 5 citations
Builds on11
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang et al.ICLR 2020 · 532 citations
- GraphDF: A Discrete Flow Model for Molecular Graph GenerationYouzhi Luo, Keqiang Yan, Shuiwang JiICML 2021 · 264 citations
- Improved Contrastive Divergence Training of Energy-Based ModelsYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2021 · 171 citations
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam et al.ICLR 2020 · 147 citations
Related papers
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
- Moment Matching Denoising Gibbs SamplingMingtian Zhang, Alex Hawkins-Hooker, Brooks Paige, David BarberNeurIPS 2023 · 8 citations
- Energy-based generator matching: A neural sampler for general state spaceDongyeop Woo, Minsu Kim, Minkyu Kim, Kiyoung Seong et al.NeurIPS 2025 · 3 citations
- Perturb-and-max-product: Sampling and learning in discrete energy-based modelsMiguel Lázaro-Gredilla, Antoine Dedieu, Dileep GeorgeNeurIPS 2021 · 10 citations
- Scalable Discrete Diffusion Samplers: Combinatorial Optimization and Statistical PhysicsSebastian Sanokowski, Wilhelm Franz Berghammer, Haoyu Peter Wang, Martin Ennemoser et al.ICLR 2025
