Lune

ACL2026顶会

WildReward: Learning Reward Models from In-the-Wild Human Interactions

Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao, Lei Hou, Juanzi Li

2026年份
3被引次数

摘要

Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deployment of LLMs, in-the-wild interactions have emerged as a rich source of implicit reward signals. This raises the question: Can we develop reward models directly from in-the-wild interactions? In this work, we explore this possibility by adopting WildChat as an interaction source and proposing a pipeline to extract reliable human feedback, yielding 186k high-quality instances for training WILDREWARD via ordinal regression directly on user feedback without preference pairs. Extensive experiments demonstrate that WILDREWARD achieves comparable or even superior performance compared to conventional reward models, with improved calibration and cross-sample consistency. We also observe that WILDREWARD benefits directly from user diversity, where more users yield stronger reward models. Finally, we apply WILDREWARD to online DPO training and observe significant improvements across various tasks. Code and data are released at https: //github.com/THU-KEG/WildReward .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖