RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
David Reber, Sean M. Richardson, Todd Nief, Cristina Garbacea, Victor Veitch
摘要
Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what they are actually rewarding. In this paper, we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective method for measuring the sensitivity of a reward model to high-level attributes of responses, such as sentiment, helpfulness, or complexity. Importantly, RATE measures the causal effect of an attribute on the reward. RATE uses LLMs to rewrite responses to produce imperfect counterfactual examples that can be used to measure causal effects. A key challenge is that these rewrites are imperfect in a manner that can induce substantial bias in the estimated sensitivity of the reward model to the target attribute. The core idea of RATE is to adjust for this imperfect-rewrite effect by rewriting twice. We establish the validity of the RATE procedure and show empirically that it is an effective estimator. Code is available at https://github.com/toddnief/RATE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward ModelingDengcan Liu, Fengkai Yang, Xiaohan Wang, Shurui Yan 等KDD 2026 · 被引用 12 次
- Preference Learning for AI Alignment: a Causal PerspectiveKasia Kobalczyk, Mihaela van der SchaarICML 2025
它引用的顶会 Paper7
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 被引用 293 次
- CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorEldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat 等NeurIPS 2022 · 被引用 69 次
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 被引用 19 次
- Causal Estimation for Text Data with (Apparent) Overlap ViolationsLin Gui, Victor VeitchICLR 2023 · 被引用 6 次
相关 Paper
- Interpreting Language Reward Models via Contrastive ExplanationsJunqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lécué 等ICLR 2025
- Debiasing Reward Models via Causally Motivated Inference-Time InterventionKazutoshi Shinoda, Kosuke Nishida, Kyosuke NishidaACL 2026
- Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin 等ICLR 2025
- Mitigating Length Bias in RLHF Through a Causal LensHyeonji Kim, Sujeong Oh, Sanghack LeeAAAI 2026 · 被引用 3 次
- Reward Models Inherit Value Biases from PretrainingBrian R. Christian, Jessica A. F. Thompson, Elle Michelle Yang, Vincent Adam 等ICLR 2026 · 被引用 4 次
