Interpreting Language Reward Models via Contrastive Explanations
Junqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lécué, Manuela Veloso
Abstract
Reward models (RMs) are a crucial component in the alignment of large language models' (LLMs) outputs with human values. RMs approximate human preferences over possible LLM responses to the same prompt by predicting and comparing reward scores. However, as they are typically modified versions of LLMs with scalar output heads, RMs are large black boxes whose predictions are not explainable. More transparent RMs would enable improved trust in the alignment of LLMs. In this work, we propose to use contrastive explanations to explain any binary response comparison made by an RM. Specifically, we generate a diverse set of new comparisons similar to the original one to characterise the RM's local behaviour. The perturbed responses forming the new comparisons are generated to explicitly modify manually specified high-level evaluation attributes, on which analyses of RM behaviour are grounded. In quantitative experiments, we validate the effectiveness of our method for finding high-quality contrastive explanations. We then showcase the qualitative usefulness of our method for investigating global sensitivity of RMs to each evaluation attribute, and demonstrate how representative examples can be automatically extracted to explain and compare behaviours of different RMs. We see our method as a flexible framework for RM explanation, providing a basis for more interpretable and trustworthy LLM alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce9b486e-bf64-4fbc-9969-ff1d8a96e1edCited by top-tier papers5
- Representation Consistency for Accurate and Coherent LLM Answer AggregationJunqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante et al.NeurIPS 2025 · 6 citations
- Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward ModelingPankayaraj Pathmanathan, Furong HuangACL 2026 · 2 citations
- Automatically Finding Reward Model BiasesAtticus Wang, Iván Arcuschin, Arthur ConmyICML 2026 · 1 citation
- RATE: Causal Explainability of Reward Models with Imperfect CounterfactualsDavid Reber, Sean M. Richardson, Todd Nief, Cristina Garbacea et al.ICML 2025
- Discovering Implicit Large Language Model Alignment ObjectivesEdward Chen, Sanmi Koyejo, Carlos GuestrinICML 2026
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye et al.ICML 2024 · 274 citations
Related papers
- Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in ExplanationsPedro Lobato Ferreira, Wilker Aziz, Ivan TitovICLR 2026 · 12 citations
- RMB: Comprehensively benchmarking reward models in LLM alignmentEnyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi et al.ICLR 2025
- Reward Models Inherit Value Biases from PretrainingBrian R. Christian, Jessica A. F. Thompson, Elle Michelle Yang, Vincent Adam et al.ICLR 2026 · 4 citations
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li et al.NeurIPS 2024 · 76 citations
- SparseRM: A Lightweight Preference Modeling with Sparse AutoencoderDengcan Liu, Jiahao Li, Zheren Fu, Yi Tu et al.AAAI 2026
