Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
Youliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan, Xiaoyuan Liu, Junjielong Xu, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Pinjia He
摘要
In this paper, we observe that current models are susceptible to reward hacking, leading to a substantial overestimation of a model's reasoning ability. This is evidenced by a high incidence of "false positives"-solutions that reach the correct answer through an unsound process. Through a systematic analysis with human verification, we establish a taxonomy of these failure modes, identifying patterns like Miracle Steps-abrupt jumps to a correct output without a valid preceding derivation. Probing experiments suggest that these Miracle Steps are linked to answer-recall shortcuts, including memorization from pretraining, where the model accesses the correct answer independently of its reasoning chain. To mitigate this systemic issue, we introduce the Rubric Reward Model (RRM), a process-oriented reward function that evaluates the entire reasoning trajectory against problem-specific rubrics. The RRM explicitly penalizes logical flaws and encourages rigorous deduction. When integrated into an RL pipeline, RRM-based training consistently outperforms outcome-only supervision across four math benchmarks. Notably, it boosts Verified Pass@1024 on AIME2024 from 26.7% to 62.6% and reduces the incidence of Miracle Steps by 71%. Our work demonstrates that rewarding the solution process is crucial for building accurate and reliable models. 1 * This work was completed before the author's affiliation with UC Berkeley. † Pinjia He is the corresponding author. 1 We released our code and data at https://github.com/ YouliangYuan/rrm-cure-miracle-steps .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath 等ICLR 2026 · 被引用 340 次
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux 等NeurIPS 2025 · 被引用 291 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
相关 Paper
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningJie Cheng, Gang Xiong, Ruixi Qiao, Lijun Li 等NeurIPS 2025 · 被引用 56 次
- Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningWeijie Li, Jin Wang, Liang-Chih Yu, Xuejie ZhangAAAI 2026 · 被引用 1 次
- Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward ModelsRishabh Tiwari, Aditya Tomar, Udbhav Bamba, Monishwaran Maheswaran 等ICML 2026
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin 等ICLR 2026 · 被引用 147 次
- Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor CorrectionRuike Song, Zeen Song, Huijie Guo, Wenwen QiangAAAI 2026 · 被引用 2 次
