Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
Pedro Lobato Ferreira, Wilker Aziz, Ivan Titov
Abstract
Chain-of-thought explanations are widely used to inspect the decision process of large language models (LLMs) and to evaluate the trustworthiness of model outputs, making them important for effective collaboration between LLMs and humans. We demonstrate that preference optimization -- a key step in the alignment phase -- can inadvertently reduce the faithfulness of these explanations. This occurs because the reward model (RM), which guides alignment, is tasked with optimizing both the expected quality of the response and the appropriateness of the explanations (e.g., minimizing bias or adhering to safety standards), creating potential conflicts. The RM lacks a mechanism to assess the consistency between the model’s internal decision process and the generated explanation. Consequently, the LLM may engage in ``reward hacking'' by producing a final response that scores highly while giving an explanation tailored to maximize reward rather than accurately reflecting its reasoning. To address this issue, we propose enriching the RM’s input with a causal attribution of the prediction, allowing the RM to detect discrepancies between the generated self-explanation and the model's decision process. In controlled settings, we show that this approach reduces the tendency of the LLM to generate misleading explanations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 507d9ffa-58b7-4075-9e97-12eda0600f11Cited by top-tier papers2
- Markovian Transformers for Informative Language ModelingScott Viteri, Max Lamparth, Peter Chatain, Clark W. BarrettICLR 2026 · 3 citations
- CLARity: Reasoning Consistency Alone Can Teach Reinforced ExpertsJiuheng Lin, Cong Jiang, Zirui Wu, Jiarui Sun et al.ACL 2026
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Interpreting Language Reward Models via Contrastive ExplanationsJunqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.ICLR 2025
- Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language ModelsWei Jie Yeo, Ranjan Satapathy, Erik CambriaEMNLP 2025 · 2 citations
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi et al.ICML 2024 · 145 citations
- Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic InferenceJiazheng Li, Hanqi Yan, Yulan HeACL 2025
- Scaling Laws for Reward Model Overoptimization in Direct Alignment AlgorithmsRafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi et al.NeurIPS 2024 · 169 citations
