BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
Gaurav Pandey, Yatin Nandwani, Tahira Naseem, Mayank Mishra, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo
摘要
Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
- Statistical Rejection Sampling Improves Preference OptimizationTianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman 等ICLR 2024 · 被引用 346 次
相关 Paper
- Aligning Language Models with Preferences through f-divergence MinimizationDongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen 等ICML 2023 · 被引用 119 次
- Human Alignment of Large Language Models through Online Preference OptimisationDaniele Calandriello, Zhaohan Daniel Guo, Rémi Munos, Mark Rowland 等ICML 2024 · 被引用 90 次
- Design Considerations in Offline Preference-based RLAlekh Agarwal, Christoph Dann, Teodor Vanislavov MarinovICML 2025
- GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNetsOh Joon Kwon, Daiki E. Matsunaga, Kee-Eung KimEMNLP 2024 · 被引用 5 次
- Aligning Language Models with Human Preferences via a Bayesian ApproachJiashuo Wang, Haozhao Wang, Shichao Sun, Wenjie LiNeurIPS 2023 · 被引用 42 次
