Gradient-based Constrained Sampling from Language Models
Sachin Kumar, Biswajit Paria, Yulia Tsvetkov
Abstract
Large pretrained language models generate fluent text but are notoriously hard to controllably sample from. In this work, we study constrained sampling from such language models: generating text that satisfies user-defined constraints, while maintaining fluency and model's performance in a downstream task. We propose MUCOLA-a sampling procedure that combines the log-likelihood of the language model with arbitrary (differentiable) constraints in a single energy function, and then generates samples in a non-autoregressive manner. Specifically, it initializes the entire output sequence with noise and follows a Markov chain defined by Langevin Dynamics using the gradients of the energy function. We evaluate MUCOLA on text generation with soft and hard constraints as well as their combinations obtaining significant improvements over competitive baselines for toxicity avoidance, sentiment control, and keyword-guided generation. 1 cat Transformer Layers
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32039627-ca3a-4b68-ae53-48528528ce07Cited by top-tier papers24
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum et al.NeurIPS 2023 · 454 citations
- Controlled Text Generation with Natural Language InstructionsWangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell et al.ICML 2023 · 121 citations
- Decoding-Time Language Model Alignment with Multiple ObjectivesRuizhe Shi, Yifang Chen, Yushi Hu, Alisa Liu et al.NeurIPS 2024 · 111 citations
- Grammar-Aligned DecodingKanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova et al.NeurIPS 2024 · 73 citations
- Controlled Text Generation via Language Model ArithmeticJasper Dekoninck, Marc Fischer, Luca Beurer-Kellner, Martin T. VechevICLR 2024 · 58 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Controlled LLM Decoding via Discrete Auto-regressive BiasingPatrick Pynadath, Ruqi ZhangICLR 2025
- Controlled Text Generation as Continuous Optimization with Multiple ConstraintsSachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia TsvetkovNeurIPS 2021 · 91 citations
- Principled Gradient-Based MCMC for Conditional Sampling of TextLi Du, Afra Amini, Lucas Torroba Hennigen, Xinyan Velocity Yu et al.ICML 2024 · 6 citations
- Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models DecodingLifu Tu, Semih Yavuz, Jin Qu, Jiacheng Xu et al.EMNLP 2024 · 2 citations
- Controllable Generation via Locally Constrained ResamplingKareem Ahmed, Kai-Wei Chang, Guy Van den BroeckICLR 2025
