Self-Evolved Reward Learning for LLMS
Chenghua Huang, Zhizhen Fan, Lu Wang, Fangkai Yang, Pu Zhao, Zeqi Lin, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, Qi Zhang
摘要
Reinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences, playing a pivotal role in the success of conversational models like GPT-4, ChatGPT, and Llama 2. A core challenge in employing RLHF lies in training a reliable reward model (RM), which relies on high-quality labels typically provided by human experts or advanced AI system. These methods can be costly and may introduce biases that affect the language model's responses. As language models improve, human input may become less effective in further enhancing their performance. In this paper, we propose Self-Evolved Reward Learning (SER), a novel approach where the RM generates additional training data to iteratively improve itself. We conducted extensive experiments on multiple datasets such as HH-RLHF and UltraFeedback, using models like Mistral and Llama 3, and compare SER against various baselines. Our results demonstrate that even with limited human-annotated data, learning from self-feedback can robustly enhance RM performance, thereby boosting the capabilities of large language models (LLMs). Resources of this paper can be found at https://aka.ms/ser
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Text2Grad: Reinforcement Learning from Natural Language FeedbackHanyang Wang, Lu Wang, Chaoyun Zhang, Tianjun Mao 等ICLR 2026 · 被引用 18 次
- A Survey on Efficient Large Language Model Training: From Data-centric PerspectivesJunyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao 等ACL 2025 · 被引用 12 次
- CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated RewardsZhiming Lin, Kai Zhao, Sophie Zhang, Peilai Yu 等AAAI 2026 · 被引用 11 次
- Limited Preference Data? Learning Better Reward Model with Latent Space SynthesisLeitian Tao, Xuefeng Du, Sharon LiNeurIPS 2025 · 被引用 2 次
- RLTHF: Targeted Human Feedback for LLM AlignmentYifei Xu, Tusher Chakraborty, Emre Kiciman, Bibek Aryal 等ICML 2025
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang 等NeurIPS 2023 · 被引用 515 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
相关 Paper
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen 等ACL 2026 · 被引用 2 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
- Real-Time Aligned Reward Model beyond SemanticsZixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng 等ICML 2026 · 被引用 18 次
- Provably Efficient Online RLHF with One-Pass Reward ModelingLong-Fei Li, Yu-Yang Qian, Peng Zhao, Zhi-Hua ZhouNeurIPS 2025 · 被引用 8 次
- Self-Play Preference Optimization for Language Model AlignmentYue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji 等ICLR 2025
