BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
Gaurav Pandey, Yatin Nandwani, Tahira Naseem, Mayank Mishra, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo
Abstract
Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6547e0eb-a00e-40af-927d-938a0c36f606Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- Statistical Rejection Sampling Improves Preference OptimizationTianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman et al.ICLR 2024 · 346 citations
Related papers
- Aligning Language Models with Preferences through f-divergence MinimizationDongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen et al.ICML 2023 · 119 citations
- Human Alignment of Large Language Models through Online Preference OptimisationDaniele Calandriello, Zhaohan Daniel Guo, Rémi Munos, Mark Rowland et al.ICML 2024 · 90 citations
- Design Considerations in Offline Preference-based RLAlekh Agarwal, Christoph Dann, Teodor Vanislavov MarinovICML 2025
- GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNetsOh Joon Kwon, Daiki E. Matsunaga, Kee-Eung KimEMNLP 2024 · 5 citations
- Aligning Language Models with Human Preferences via a Bayesian ApproachJiashuo Wang, Haozhao Wang, Shichao Sun, Wenjie LiNeurIPS 2023 · 42 citations
