Principled Gradient-Based MCMC for Conditional Sampling of Text
Li Du, Afra Amini, Lucas Torroba Hennigen, Xinyan Velocity Yu, Holden Lee, Jason Eisner, Ryan Cotterell
Abstract
We consider the problem of sampling text from an energy-based model. This arises, for example, when sampling text from a neural language model subject to soft constraints. Although the target distribution is discrete, the internal computations of the energy function (given by the language model) are differentiable, so one would like to exploit gradient information within a method such as MCMC. Alas, all previous attempts to generalize gradient-based MCMC to text sampling fail to sample correctly from the target distribution. We propose a solution, along with variants, and study its theoretical properties. Through experiments on various forms of text generation, we demonstrate that our unbiased samplers are able to generate more fluent text while better adhering to the control objectives. The same methods could be used to sample from discrete energy-based models unrelated to text. Introduction Recent papers have performed controlled text generation from pretrained language models by formulating energybased models over text and applying Markov Chain Monte Carlo (MCMC) algorithms
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c5583ef-a847-4d32-b077-21428794dfe9Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- COLD Decoding: Energy-based Constrained Text Generation with Langevin DynamicsLianhui Qin, Sean Welleck, Daniel Khashabi, Yejin ChoiNeurIPS 2022 · 217 citations
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam et al.ICLR 2020 · 147 citations
- Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsWill Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud et al.ICML 2021 · 113 citations
Related papers
- Structured Voronoi SamplingAfra Amini, Li Du, Ryan CotterellNeurIPS 2023 · 5 citations
- Controlled LLM Decoding via Discrete Auto-regressive BiasingPatrick Pynadath, Ruqi ZhangICLR 2025
- A Distributional Approach to Controlled Text GenerationMuhammad Khalifa, Hady Elsahar, Marc DymetmanICLR 2021 · 135 citations
- Gradient-based Constrained Sampling from Language ModelsSachin Kumar, Biswajit Paria, Yulia TsvetkovEMNLP 2022 · 22 citations
- Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--HastingsKartik Goyal, Chris Dyer, Taylor Berg-KirkpatrickICLR 2022 · 53 citations
