SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
Yuyang Ding, Xinyu Shi, Juntao Li, Xiaobo Liang, Zhaopeng Tu, Min Zhang
摘要
Process reward models (PRMs) offer fine-grained, step-level evaluations that facilitate deeper reasoning processes in large language models (LLMs), proving effective in complex tasks like mathematical reasoning. However, developing PRMs is challenging due to the high cost and limited scalability of human-annotated data. Synthetic data from Monte Carlo (MC) estimation is a promising alternative but suffers from a high noise ratio, which can cause overfitting and hinder large-scale training. In this work, we conduct a preliminary study on the noise distribution in synthetic data from MC estimation, identifying that annotation models tend to both underestimate and overestimate step correctness due to limitations in their annotation capabilities. Building on these insights, we propose Self-Denoising Monte Carlo Annotation (SCAN), an efficient data synthesis and noise-tolerant learning framework. Our key findings indicate that: (1) Even lightweight models (e.g., 1.5B parameters) can produce high-quality annotations through a self-denoising strategy, enabling PRMs to achieve superior performance with only 6% the inference cost required by vanilla MC estimation. (2) With our robust learning strategy, PRMs can effectively learn from this weak supervision, achieving a 39.2 F1 score improvement (from 19.9 to 59.1) in ProcessBench. Despite using only a compact synthetic dataset, our models surpass strong baselines, including those trained on large-scale human-annotated datasets such as PRM800K. Furthermore, performance continues to improve as we scale up the synthetic data, highlighting the potential of SCAN for scalable, cost-efficient, and robust PRM training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable ReasoningYuyang Ding, Chi Zhang, Juntao Li, Haibin Lin 等ICLR 2026 · 被引用 7 次
- A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and UsageCongmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen 等ACL 2026
- Training Data Efficiency in Multimodal Process Reward ModelsJinyuan Li, Chengsong Huang, Langlin Huang, Shaoyang Xu 等ICML 2026
它引用的顶会 Paper14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
相关 Paper
- R-PRM: Reasoning-Driven Process Reward ModelingShuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen 等EMNLP 2025
- An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical ReasoningWei Sun, Qianlong Du, Fuwei Cui, Jiajun ZhangACL 2025 · 被引用 15 次
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 被引用 3 次
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsMingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou 等ACL 2025 · 被引用 85 次
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou 等AAAI 2026 · 被引用 68 次
