Quantifying Differences in Reward Functions
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, Jan Leike
Abstract
For many tasks, the reward function is inaccessible to introspection or too complex to be specified procedurally, and must instead be learned from user data. Prior work has evaluated learned reward functions by evaluating policies optimized for the learned reward. However, this method cannot distinguish between the learned reward function failing to reflect user preferences and the policy optimization process failing to optimize the learned reward. Moreover, this method can only tell us about behavior in the evaluation environment, but the reward may incentivize very different behavior in even a slightly different deployment environment. To address these problems, we introduce the Equivalent-Policy Invariant Comparison (EPIC) distance to quantify the difference between two reward functions directly, without a policy optimization step. We prove EPIC is invariant on an equivalence class of reward functions that always induce the same optimal policy. Furthermore, we find EPIC can be efficiently approximated and is more robust than baselines to the choice of coverage distribution. Finally, we show that EPIC distance bounds the regret of optimal policies even under different transition dynamics, and we confirm empirically that it predicts policy training success. Our source code is available at https: //github.com/HumanCompatibleAI/evaluating-rewards . * Work partially conducted while at DeepMind.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez et al.ICLR 2024 · 154 citations
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 60 citations
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 48 citations
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 30 citations
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer et al.ICLR 2024 · 22 citations
Related papers
- Dynamics-Aware Comparison of Learned Reward FunctionsBlake Wulfe, Logan Michael Ellis, Jean Mercat, Rowan Thomas McAllister et al.ICLR 2022 · 18 citations
- STARC: A General Framework For Quantifying Differences Between Reward FunctionsJoar Max Viktor Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner et al.ICLR 2024 · 14 citations
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low RegretLukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré et al.ICML 2025
- Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin et al.ICLR 2025
