Amortizing intractable inference in large language models
Edward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, Nikolay Malkin
Abstract
Autoregressive large language models (LLMs) compress knowledge from their training data through next-token conditional distributions. This limits tractable querying of this knowledge to start-to-end autoregressive sampling. However, many tasks of interest-including sequence continuation, infilling, and other forms of constrained generation-involve sampling from intractable posterior distributions. We address this limitation by using amortized Bayesian inference to sample from these intractable posteriors. Such amortization is algorithmically achieved by finetuning LLMs via diversity-seeking reinforcement learning algorithms: generative flow networks (GFlowNets). We empirically demonstrate that this distributionmatching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization. As an important application, we interpret chain-of-thought reasoning as a latent variable modeling problem and demonstrate that our approach enables data-efficient adaptation of LLMs to tasks that require multi-step rationalization and tool use. Code: https://github.com/GFNOrg/gfn-lm-tuning . * Equal contribution. ∞ Work done during internship at Mila. † CIFAR AI Chair. ⋄ CIFAR Senior Fellow.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdcf3d22-761c-4b14-8b48-9cccd168a74bCited by top-tier papers54
- Privileged Information Distillation for Language ModelsEmiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste et al.ICML 2026 · 61 citations
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 61 citations
- Adaptable Logical Control for Large Language ModelsHonghua Zhang, Po-Nien Kung, Masahiro Yoshida, Guy Van den Broeck et al.NeurIPS 2024 · 42 citations
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li et al.ICLR 2026 · 41 citations
- Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingBrian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain et al.NeurIPS 2025 · 34 citations
Builds on32
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
Related papers
- Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal ExamplesFangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao et al.ICML 2025
- GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksHaoqiang Kang, Enna Sachdeva, Piyush Gupta, Sangjae Bae et al.CVPR 2025
- Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVRDohyung Kim, Minbeom Kim, Jeonghye Kim, Lee Sangmook et al.ICML 2026
- PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution MatchingRuishuo Chen, Yu Chen, Zhuoran Li, Longbo HuangICML 2026 · 1 citation
- How reinforcement learning after next-token prediction facilitates learningNikolaos Tsilivis, Eran Malach, Karen Ullrich, Julia KempeICLR 2026 · 9 citations
