Is it Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He
摘要
Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior may be explicit, i.e. verbalized in the model's chain-ofthought (CoT), or implicit, where the CoT appears benign thus bypasses CoT monitors. To detect implicit reward hacking, we propose TRACE (Truncated Reasoning AUC Evaluation). Our key observation is that hacking occurs when exploiting the loophole is easier than solving the actual task. This means that the model is using less "effort" than required to achieve high reward. TRACE quantifies effort by measuring how early a model's reasoning becomes sufficient to obtain the reward. We progressively truncate a model's CoT at various lengths, force the model to answer, and estimate the expected reward at each cutoff. A hacking model, which takes a shortcut, will achieve a high expected reward with only a small fraction of its CoT, yielding a large area under the reward-vs-length curve. TRACE achieves over 65% gains over our strongest 72B CoT monitor in math reasoning, and over 30% gains over a 32B monitor in coding. We further show that TRACE can discover unknown loopholes during training. Overall, TRACE offers a scalable unsupervised approach for oversight where current monitoring methods prove ineffective. 48. A student has 7 reference books, including 2 Chinese books, […]. Calculate the total number of different ways the books can be arranged.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-ThoughtSiddharth Boppana, Annabel Ma, Max Loeffler, Raphaël Sarfati 等ICML 2026 · 被引用 30 次
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important ReasoningQinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang 等ICLR 2026 · 被引用 406 次
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan 等ICML 2026 · 被引用 175 次
- Describing Differences between Text Distributions with Natural LanguageRuiqi Zhong, Charlie Snell, Dan Klein, Jacob SteinhardtICML 2022 · 被引用 61 次
相关 Paper
- Large language models can learn and generalize steganographic chain-of-thought under process supervisionRobert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev 等NeurIPS 2025 · 被引用 29 次
- Benchmarking Reward Hack Detection in Code Environments via Contrastive AnalysisDarshan Deshpande, Anand Kannappan, Rebecca QianICML 2026 · 被引用 13 次
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool UseKunvar ThamanICML 2026 · 被引用 14 次
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky 等NeurIPS 2025 · 被引用 50 次
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric RewardsYouliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan 等ACL 2026 · 被引用 5 次
