Generative Verifiers: Reward Modeling as Next-Token Prediction
Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, Rishabh Agarwal
Abstract
Verifiers or reward models are often used to enhance the reasoning performance of large language models (LLMs). A common approach is the Best-of-N method, where N candidate solutions generated by the LLM are ranked by a verifier, and the best one is selected. While LLM-based verifiers are typically trained as discriminative classifiers to score solutions, they do not utilize the text generation capabilities of pretrained LLMs. To overcome this limitation, we instead propose training verifiers using the ubiquitous next-token prediction objective, jointly on verification and solution generation. Compared to standard verifiers, such generative verifiers (GenRM) can benefit from several advantages of LLMs: they integrate seamlessly with instruction tuning, enable chain-of-thought reasoning, and can utilize additional test-time compute via majority voting for better verification. We demonstrate that GenRM outperforms discriminative, DPO verifiers, and LLM-as-a-Judge, resulting in large performance gains with Best-of-N, namely 5% → 45.3% on algorithmic tasks and 73% → 93.4% on GSM8K. In easy-to-hard generalization settings, we observe improvements of 28% → 44.6% on MATH, and 37.9% → 53.5% on MMLU abstract algebra. Furthermore, we find that training GenRM with synthetic verification rationales is sufficient to pick out subtle errors on math problems. Finally, we demonstrate that GenRM scales favorably with model size and test-time compute.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27919433-e1f1-4682-908e-5b1a9992fc33Cited by top-tier papers170
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin et al.ICLR 2026 · 147 citations
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison et al.ICLR 2026 · 139 citations
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen et al.ICLR 2026 · 110 citations
- The Art of Scaling Reinforcement Learning Compute for LLMsDevvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal et al.ICLR 2026 · 95 citations
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong et al.ACL 2026 · 75 citations
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen et al.ACL 2026 · 3 citations
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
- Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsYi Su, Dian Yu, Linfeng Song, Juntao Li et al.ACL 2026
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen et al.ICML 2025
- RL Tango: Reinforcing Generator and Verifier Together for Language ReasoningKaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong et al.NeurIPS 2025 · 44 citations
