Generative Verifiers: Reward Modeling as Next-Token Prediction
Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, Rishabh Agarwal
摘要
Verifiers or reward models are often used to enhance the reasoning performance of large language models (LLMs). A common approach is the Best-of-N method, where N candidate solutions generated by the LLM are ranked by a verifier, and the best one is selected. While LLM-based verifiers are typically trained as discriminative classifiers to score solutions, they do not utilize the text generation capabilities of pretrained LLMs. To overcome this limitation, we instead propose training verifiers using the ubiquitous next-token prediction objective, jointly on verification and solution generation. Compared to standard verifiers, such generative verifiers (GenRM) can benefit from several advantages of LLMs: they integrate seamlessly with instruction tuning, enable chain-of-thought reasoning, and can utilize additional test-time compute via majority voting for better verification. We demonstrate that GenRM outperforms discriminative, DPO verifiers, and LLM-as-a-Judge, resulting in large performance gains with Best-of-N, namely 5% → 45.3% on algorithmic tasks and 73% → 93.4% on GSM8K. In easy-to-hard generalization settings, we observe improvements of 28% → 44.6% on MATH, and 37.9% → 53.5% on MMLU abstract algebra. Furthermore, we find that training GenRM with synthetic verification rationales is sufficient to pick out subtle errors on math problems. Finally, we demonstrate that GenRM scales favorably with model size and test-time compute.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper170
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin 等ICLR 2026 · 被引用 147 次
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison 等ICLR 2026 · 被引用 139 次
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen 等ICLR 2026 · 被引用 110 次
- The Art of Scaling Reinforcement Learning Compute for LLMsDevvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal 等ICLR 2026 · 被引用 95 次
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong 等ACL 2026 · 被引用 75 次
它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen 等ACL 2026 · 被引用 3 次
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
- Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsYi Su, Dian Yu, Linfeng Song, Juntao Li 等ACL 2026
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen 等ICML 2025
- RL Tango: Reinforcing Generator and Verifier Together for Language ReasoningKaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong 等NeurIPS 2025 · 被引用 44 次
