Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
Rishabh Tiwari, Aditya Tomar, Udbhav Bamba, Monishwaran Maheswaran, Heng Yang, Michael Mahoney, Kurt Keutzer, Amir Gholaminejad
Abstract
Process Reward Models (PRMs) are rapidly becoming the backbone of LLM reasoning pipelines, yet we demonstrate that state-of-the-art PRMs are systematically exploitable under adversarial optimization pressure. To address this, we introduce a three-tiered diagnostic framework that applies increasing adversarial pressure to quantify these vulnerabilities. Static perturbation analysis uncovers a fluency-logic dissociation: high invariance to surface-level style changes (reward changes <0.1), yet inconsistent detection of logicallycorrupted reasoning, with different models failing on different attack types. Adversarial optimization demonstrates that gradient-based attacks inflate rewards on invalid trajectories, with reward landscapes exhibiting wide, exploitable peaks. RL-induced reward hacking exposes the critical failure mode: policies trained on AIME problems achieve near-perfect PRM rewards (>0.9), while ground-truth accuracy remains low (below 4%), with 43% of reward gains attributable to stylistic shortcuts. These findings reveal that current PRMs function as fluency detectors rather than reasoning verifiers, creating systematic blind spots that undermine their use as training signals. We release PRM-BiasBench and a diagnostic toolkit to enable robustness evaluation before deployment. The code and dataset are available at https://github.com/SqueezeAILab/rewardunder-attack
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7c7080d-d788-4d0b-9b1b-930e211df7faBuilds on8
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
Related papers
- Adversarial Training for Process Reward ModelsGurusha Juneja, Deepak Nathani, William WangICML 2026 · 2 citations
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsMingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou et al.ACL 2025 · 85 citations
- Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language ModelsYuansen Liu, Yixuan Tang, Anthony Kum Hoe TungACL 2026
- rePIRL: Learn PRM with Inverse RL for LLM ReasoningXian Wu, Kaijie Zhu, Ying Zhang, Lun Wang et al.ICML 2026
- Unbiased Principles, Robust RewardsQingnan Ren, Zhen Fang, Shiting Huang, Yu Zeng et al.ICML 2026
