Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
Ruike Song, Zeen Song, Huijie Guo, Wenwen Qiang
Abstract
External reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving. However, these systems are prone to reward hacking, where high-scoring but logically incorrect paths are assigned high scores by the PRMs, leading to incorrect answers. From a causal inference perspective, we attribute this phenomenon primarily to the presence of confounding semantic features. To address it, we propose Causal Reward Adjustment (CRA), a method that mitigates reward hacking by estimating the true reward of a reasoning path. CRA trains sparse autoencoders on the PRM’s internal activations to recover interpretable features, then corrects confounding by using backdoor adjustment. Experiments on math solving datasets demonstrate that CRA mitigates reward hacking and improves final accuracy, without modifying the policy model or retraining PRM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
Related papers
- Linking Process to Outcome: Conditional Reward Modeling for LLM ReasoningZheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li et al.ICLR 2026 · 16 citations
- Factored Causal Representation Learning for Robust Reward Modeling in RLHFYupei Yang, Lin Yang, Wanxi Deng, Lin Qu et al.ICML 2026 · 1 citation
- Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in ExplanationsPedro Lobato Ferreira, Wilker Aziz, Ivan TitovICLR 2026 · 12 citations
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningJie Cheng, Gang Xiong, Ruixi Qiao, Lijun Li et al.NeurIPS 2025 · 56 citations
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric RewardsYouliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan et al.ACL 2026 · 5 citations
