Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection
Sadegh Mahdavi, Branislav Kisacanin, Shubham Toshniwal, Wei Du, Ivan Moshkov, George Armstrong, Renjie Liao, Christos Thrampoulidis, Igor Gitman
摘要
Large language models have achieved remarkable success on final-answer mathematical problems, largely due to the ease of applying reinforcement learning with verifiable rewards. However, the reasoning underlying these solutions is often flawed. Advancing to rigorous proof-based mathematics requires reliable proof verification capabilities. We begin by analyzing multiple evaluation setups and show that focusing on a single benchmark can lead to brittle or misleading conclusions. To address this, we evaluate both proof-based and final-answer reasoning to obtain a more reliable measure of model performance. We then scale two major generative verification methods (GenSelect and LLM-as-a-Judge) to millions of tokens and identify their combination as the most effective framework for solution verification and selection. We further show that the choice of prompt for LLM-as-a-Judge significantly affects the model's performance, but reinforcement learning can reduce this sensitivity. However, despite improving proof-level metrics, reinforcement learning does not enhance final-answer precision, indicating that current models often reward stylistic or procedural correctness rather than mathematical validity. Our results establish practical guidelines for designing and evaluating scalable proof-verification and selection systems. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- V1: Unifying Generation and Self-Verification for Parallel ReasonersHarman Singh, Xiuyu Li, Kusha Sareen, Monishwaran Maheswaran 等ICML 2026 · 被引用 8 次
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement LearningMing Chen, Sheng Tang, Rong-Xi Tan, Ziniu Li 等ICML 2026 · 被引用 2 次
- Pessimistic Verification for Open-Ended Math QuestionsYanxing Huang, Zihan Tang, Zejin Lin, Peng Li 等ICML 2026 · 被引用 2 次
- QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical ProofsSantiago Gonzalez, Alireza Amiribavandpour, Peter Ye, Edward Zhang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper10
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
相关 Paper
- Reliable Fine-Grained Evaluation of Natural Language Math ProofsWenjie Ma, Andrei Cojocaru, Neel Kolhe, Haihan Zhang 等ICLR 2026 · 被引用 14 次
- Reinforcing General Reasoning Without VerifiersXiangxin Zhou, Zichen Liu, Anya Sims, Haonan Wang 等ICLR 2026 · 被引用 75 次
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo 等AAAI 2026 · 被引用 16 次
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi 等ICLR 2025
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
