PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
Xiangfeng Wang, Hangyu Guo, Yanlin Lai, Mitt Huang, Liang Zhao, Chengyuan Yao, Yinmin Zhang, Qi Han, Xiaoxiao Ren, Chun Yuan, Tong Xu, Zheng Ge
摘要
While model-based verifiers are essential for scaling Reinforcement Learning with Verifiable Rewards (RLVR), current outcome-centric verification paradigms primarily focus on the consistency between the final result and the ground truth, often neglecting potential errors in the derivation process. This leads to assigning positive rewards to correct answers produced from incorrect derivations. To bridge this gap, we introduce P R I M E, a benchmark for evaluating verifiers on PRocess-outcome alignment verification In Mathematics and Engineering. Curated from a comprehensive collection of college-level STEM problems, P R I M E comprises 2,530 high-difficulty samples through a consistency-based filtering pipeline. Through extensive evaluation, we find that current verifiers frequently fail to detect derivation flaws. Furthermore, we propose a process-aware RLVR training paradigm utilizing verifiers selected via P R I M E. This approach substantially outperforms the outcome-only verification baseline, achieving absolute performance gains of 8.29%, 9.12%, and 7.31% on AIME24, AIME25, and Beyond-AIME, respectively, for the Qwen3-14B-Base model. Finally, we demonstrate a strong linear correlation (𝑅 2 > 0.92) between verifier accuracy on P R I M E and RLVR training effectiveness, validating P R I M E as a reliable predictor for verifier selection. We release the benchmark and code at https://github.com/wonderful9462/PRIME .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SCI-Verifier: Scientific Verifier with ThinkingShenghe Zheng, Chenyu Huang, Fangchen Yu, Junchi Yao 等ICLR 2026 · 被引用 5 次
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsPeiyi Wang, Lei Li, Zhihong Shao, Runxin Xu 等ACL 2024
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi 等ICLR 2025
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao 等EMNLP 2025
相关 Paper
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo 等AAAI 2026 · 被引用 16 次
- Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And SelectionSadegh Mahdavi, Branislav Kisacanin, Shubham Toshniwal, Wei Du 等ICML 2026 · 被引用 10 次
- Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsYi Su, Dian Yu, Linfeng Song, Juntao Li 等ACL 2026
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen 等ACL 2026 · 被引用 3 次
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathShrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming 等ACL 2026 · 被引用 13 次
