Quantifying Differences in Reward Functions
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, Jan Leike
摘要
For many tasks, the reward function is inaccessible to introspection or too complex to be specified procedurally, and must instead be learned from user data. Prior work has evaluated learned reward functions by evaluating policies optimized for the learned reward. However, this method cannot distinguish between the learned reward function failing to reflect user preferences and the policy optimization process failing to optimize the learned reward. Moreover, this method can only tell us about behavior in the evaluation environment, but the reward may incentivize very different behavior in even a slightly different deployment environment. To address these problems, we introduce the Equivalent-Policy Invariant Comparison (EPIC) distance to quantify the difference between two reward functions directly, without a policy optimization step. We prove EPIC is invariant on an equivalence class of reward functions that always induce the same optimal policy. Furthermore, we find EPIC can be efficiently approximated and is more robust than baselines to the choice of coverage distribution. Finally, we show that EPIC distance bounds the regret of optimal policies even under different transition dynamics, and we confirm empirically that it predicts policy training success. Our source code is available at https: //github.com/HumanCompatibleAI/evaluating-rewards . * Work partially conducted while at DeepMind.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez 等ICLR 2024 · 被引用 154 次
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 被引用 60 次
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 被引用 48 次
- Inverse Decision Modeling: Learning Interpretable Representations of BehaviorDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarICML 2021 · 被引用 30 次
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer 等ICLR 2024 · 被引用 22 次
相关 Paper
- Dynamics-Aware Comparison of Learned Reward FunctionsBlake Wulfe, Logan Michael Ellis, Jean Mercat, Rowan Thomas McAllister 等ICLR 2022 · 被引用 18 次
- STARC: A General Framework For Quantifying Differences Between Reward FunctionsJoar Max Viktor Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner 等ICLR 2024 · 被引用 14 次
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 被引用 92 次
- The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low RegretLukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré 等ICML 2025
- Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin 等ICLR 2025
