STARC: A General Framework For Quantifying Differences Between Reward Functions
Joar Max Viktor Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner, Adam Gleave, Alessandro Abate
Abstract
In order to solve a task using reinforcement learning, it is necessary to first formalise the goal of that task as a reward function. However, for many real-world tasks, it is very difficult to manually specify a reward function that never incentivises undesirable behaviour. As a result, it is increasingly popular to use reward learning algorithms, which attempt to learn a reward function from data. However, the theoretical foundations of reward learning are not yet well-developed. In particular, it is typically not known when a given reward learning algorithm with high probability will learn a reward function that is safe to optimise. This means that reward learning algorithms generally must be evaluated empirically, which is expensive, and that their failure modes are difficult to anticipate in advance. One of the roadblocks to deriving better theoretical guarantees is the lack of good methods for quantifying the difference between reward functions. In this paper we provide a solution to this problem, in the form of a class of pseudometrics on the space of all reward functions that we call STARC (STAndardised Reward Comparison) metrics. We show that STARC metrics induce both an upper and a lower bound on worst-case regret, which implies that our metrics are tight, and that any metric with the same properties must be bilipschitz equivalent to ours. Moreover, we also identify a number of issues with reward metrics proposed by earlier works. Finally, we evaluate our metrics empirically, to demonstrate their practical efficacy. STARC metrics can be used to make both theoretical and empirical analysis of reward learning algorithms both easier and more principled.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer et al.ICLR 2024 · 22 citations
- Quantifying the Sensitivity of Inverse Reinforcement Learning to MisspecificationJoar Max Viktor Skalse, Alessandro AbateICLR 2024 · 5 citations
- Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement LearningJunseok Kim, Dohyeong Kim, Mineui Hong, Songhwai OhICML 2026 · 1 citation
- Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin et al.ICLR 2025
- The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low RegretLukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré et al.ICML 2025
Builds on4
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- Quantifying Differences in Reward FunctionsAdam Gleave, Michael Dennis, Shane Legg, Stuart Russell et al.ICLR 2021 · 77 citations
- Dynamics-Aware Comparison of Learned Reward FunctionsBlake Wulfe, Logan Michael Ellis, Jean Mercat, Rowan Thomas McAllister et al.ICLR 2022 · 18 citations
Related papers
- Reward Learning through Ranking Mean Squared ErrorChaitanya Kharyal, Calarina Muslimani, Matthew TaylorICML 2026
- Offline Reinforcement Learning with Pseudometric LearningRobert Dadashi, Shideh Rezaeifar, Nino Vieillard, Léonard Hussenot et al.ICML 2021 · 44 citations
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum et al.AAAI 2023 · 103 citations
- Reward Model Evaluation via Automatically-Ranked Policy AlignmentAoran Wang, Lei Ou, Yang Yu, Zongzhang ZhangAAAI 2026
- Evaluating the Performance of Reinforcement Learning AlgorithmsScott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang et al.ICML 2020 · 59 citations
