Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback
Wei Shen, Guanlin Liu, Yu Yue, Ruofei Zhu, Qingping Yang, Chao Xin, Lin Yan
Abstract
Reinforcement Learning from Human Feedback (RLHF) is essential for aligning large language models (LLMs) with human preferences and values. While recent research has primarily focused on algorithmic advancements-such as reducing computational overhead or strengthening reward models to mitigate reward hacking-the critical role of prompt-data construction and its scalability has received comparatively less attention. In this paper, we address this gap by systematically exploring data-driven bottlenecks that currently hinder RLHF performance scaling, focusing specifically on the challenges posed by reward hacking and decreasing response diversity. To mitigate reward hacking, we introduce a hybrid reward system combining reasoning task verifiers (RTV) and a generative reward model (GenRM). This approach enables accurate assessment of responses against clearly defined ground-truth solutions. Additionally, in order to ensure response diversity and enhance learning effectiveness, we propose a novel prompt-selection method named Pre-PPO, explicitly identifying training prompts that are inherently challenging and thus less prone to reward hacking. Furthermore, we find that prioritizing mathematical and coding tasks during the early phases of RLHF training significantly boosts performance, given that these tasks naturally encode fine-grained response distinctions and possess clearly defined ground truths. Through experiments conducted on both small and large models, we demonstrate the effectiveness and scalability of our proposed methods. Our approach exhibits robust generalization capabilities, especially on challenging and out-of-distribution tasks, while yielding significant improvements in mathematics-intensive (STEM) and coding domains. Moreover, the proposed strategies enable the model to effectively capture subtle, task-specific distinctions during the RLHF process, substantially enhancing overall model performance. This work emphasizes the critical role of careful data construction and provides practical methodologies for addressing key performance bottlenecks in RLHF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d52cbc4-1907-4d05-b58b-e39b8d42b4e6Cited by top-tier papers8
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesJiangjie Chen, Qianyu He, Siyu Yuan, Aili Chen et al.NeurIPS 2025 · 60 citations
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVRXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen et al.ICLR 2026 · 57 citations
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM ReasoningXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yang Wang et al.NeurIPS 2025 · 41 citations
- Real-Time Aligned Reward Model beyond SemanticsZixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng et al.ICML 2026 · 18 citations
- G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior SimulationBoyu Chen, Siran Chen, Zhengrong Yue, Kainan Yan et al.AAAI 2026 · 7 citations
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina et al.ICLR 2024 · 332 citations
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 208 citations
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin et al.ICML 2024 · 165 citations
Related papers
- ODIN: Disentangled Reward Mitigates Hacking in RLHFLichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia et al.ICML 2024 · 119 citations
- Reinforcement Learning on Pre-Training DataSiheng Li, Kejiao Li, Zenan Xu, Guanhua Huang et al.ACL 2026 · 11 citations
- Factored Causal Representation Learning for Robust Reward Modeling in RLHFYupei Yang, Lin Yang, Wanxi Deng, Lin Qu et al.ICML 2026 · 1 citation
- Prototypical Reward Network for Data-Efficient RLHFJinghan Zhang, Xiting Wang, Yiqiao Jin, Changyu Chen et al.ACL 2024
- On Teacher Hacking in Language Model DistillationDaniil Tiapkin, Daniele Calandriello, Johan Ferret, Sarah Perrin et al.ICML 2025
