Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment
Yunyi Shen, Hao Sun, Jean-Francois Ton
摘要
Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research. Given the scarcity and high cost of human annotation, how to select the most informative pairs to annotate is an essential yet challenging open problem. In this work, we highlight the insight that an ideal comparison dataset for reward modeling should balance exploration of the representation space and make informative comparisons between pairs with moderate reward differences. Technically, challenges arise in quantifying the two objectives and efficiently prioritizing the comparisons to be annotated. To address this, we propose the Fisher information-based selection strategies, adapt theories from the classical experimental design literature, and apply them to the final linear layer of the deep neural network-based reward modeling tasks. Empirically, our method demonstrates remarkable performance, high computational efficiency, and stability compared to other selection methods from deep learning and classical statistical literature across multiple opensource LLMs and datasets. Further ablation studies reveal that incorporating cross-prompt comparisons in active reward modeling significantly enhances labeling efficiency, shedding light on the potential for improved annotation strategies in RLHF. Code and embeddings to reproduce all results of this paper are available at https: //github.com/YunyiShen/ARM-FI/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Towards Understanding Valuable Preference Data for Large Language Model AlignmentZizhuo Zhang, Qizhou Wang, Shanshan Ye, Jianing Zhu 等ICLR 2026 · 被引用 6 次
- CoAct: Co-Active LLM Preference Learning with Human-AI SynergyRuiyao Xu, Mihir Parmar, Tiankai Yang, Zhengyu Hu 等ACL 2026 · 被引用 1 次
- What Does Preference Learning Recover from Pairwise Comparison Data?Rattana Pukdee, Nina Balcan, Pradeep RavikumarICML 2026 · 被引用 1 次
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward ModelsTong Xie, Ching-Yuan Bai, Yuanhao Ban, Yunqi Hong 等ICML 2026
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Statistical Rejection Sampling Improves Preference OptimizationTianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman 等ICLR 2024 · 被引用 346 次
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 被引用 208 次
- Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference AdjustmentRui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu 等ICML 2024 · 被引用 144 次
相关 Paper
- RLTHF: Targeted Human Feedback for LLM AlignmentYifei Xu, Tusher Chakraborty, Emre Kiciman, Bibek Aryal 等ICML 2025
- Avoiding exp(R) scaling in RLHF through Preference-based ExplorationMingyu Chen, Yiding Chen, Wen Sun, Xuezhou ZhangNeurIPS 2025 · 被引用 9 次
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen 等ICML 2026 · 被引用 2 次
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient AlignmentXiaoqiang Lin, Arun Verma, Zhongxiang Dai, Daniela Rus 等ICLR 2026 · 被引用 12 次
- Influence-based Online Experience Selection for Effective RLHFYifan Gong, Jing Yao, Xiting Wang, Xunlong Wang 等ACL 2026
