Lune

ICLR2024顶会

The Trickle-down Impact of Reward Inconsistency on RLHF

Lingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu

2024年份
10被引次数
17顶会引用

摘要

A standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generation. A notable subject that is understudied is the (in-)consistency of RMs -whether they can recognize the semantic changes to different prompts and appropriately adapt its reward assignments -and its impact on the downstream RLHF model. In this paper, we visit a series of research questions relevant to RM inconsistency:

(1) How can we measure the consistency of reward models? (2) How consistent are the existing RMs and how can we improve them? ( 3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training? We propose CONTRAST INSTRUCTIONS -a benchmarking strategy for the consistency of RM. Each example in CONTRAST INSTRUCTIONS features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on CONTRAST INSTRUCTIONS compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques CONVEXDA and REWARDFU-SION, which enhance reward consistency through extrapolation during the RM training and inference stage, respectively. We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process.

  • Most of the work done while Lingfeng and Sihao were interns at the Tencent AI Lab.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper17

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖