RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, Juanzi Li
Abstract
Reward models are critical in techniques like Reinforcement Learning from Human Feedback (RLHF) and Inference Scaling Laws, where they guide language model alignment and select optimal responses. Despite their importance, existing reward model benchmarks often evaluate models by asking them to distinguish between responses generated by models of varying power. However, this approach fails to assess reward models on subtle but critical content changes and variations in style, resulting in a low correlation with policy model performance. To this end, we introduce RM-BENCH, a novel benchmark designed to evaluate reward models based on their sensitivity to subtle content differences and resistance to style biases. Extensive experiments demonstrate that RM-BENCH strongly correlates with policy model performance, making it a reliable reference for selecting reward models to align language models effectively. We evaluate nearly 40 reward models on RM-BENCH. Our results reveal that even state-of-the-art models achieve an average performance of only 46.6%, which falls short of random-level accuracy (50%) when faced with style bias interference. These findings highlight the significant room for improvement in current reward models. Related code and data are available at https://github.com/THU-KEG/RM-Bench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da74f544-1de4-43b6-803d-2e3012af68ceCited by top-tier papers66
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin et al.ICLR 2026 · 147 citations
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison et al.ICLR 2026 · 139 citations
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen et al.ICLR 2026 · 110 citations
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityJiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia et al.ICML 2026 · 102 citations
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
Related papers
- M-RewardBench: Evaluating Reward Models in Multilingual SettingsSrishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary et al.ACL 2025
- reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed InputsZhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Yoon Kim et al.EMNLP 2025
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsChenglong Wang, Yifu Huo, Yang Gan, Yongyu Mu et al.AAAI 2026 · 1 citation
- RMB: Comprehensively benchmarking reward models in LLM alignmentEnyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi et al.ICLR 2025
- Rethinking Reward Model Evaluation Through the Lens of Reward OveroptimizationSunghwan Kim, Dongjin Kang, Taeyoon Kwon, Hyungjoo Chae et al.ACL 2025
