The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task Specifications
Serena Booth, W. Bradley Knox, Julie Shah, Scott Niekum, Peter Stone, Alessandro Allievi
Abstract
In reinforcement learning (RL), a reward function that aligns exactly with a task's true performance metric is often sparse. For example, a true task metric might encode a reward of 1 upon success and 0 otherwise. These sparse task metrics can be hard to learn from, so in practice they are often replaced with alternative dense reward functions. These dense reward functions are typically designed by experts through an ad hoc process of trial and error. In this process, experts manually search for a reward function that improves performance with respect to the task metric while also enabling an RL algorithm to learn faster. One question this process raises is whether the same reward function is optimal for all algorithms, or, put differently, whether the reward function can be overfit to a particular algorithm. In this paper, we study the consequences of this wide yet unexamined practice of trial-and-error reward design. We first conduct computational experiments that confirm that reward functions can be overfit to learning algorithms and their hyperparameters. To broadly examine ad hoc reward design, we also conduct a controlled observation study which emulates expert practitioners' typical reward design experiences. Here, we similarly find evidence of reward function overfitting. We also find that experts' typical approach to reward design-of adopting a myopic strategy and weighing the relative goodness of each state-action pair-leads to misdesign through invalid task specifications, since RL algorithms use cumulative reward rather than rewards for individual state-action pairs as an optimization target. Code, data: github.com/serenabooth/reward-design-perils.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f9c2fd9-19fa-4cca-bae9-3c77bd1ee7ebCited by top-tier papers26
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang et al.ICLR 2024 · 582 citations
- SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningHaozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang et al.ICLR 2026 · 170 citations
- A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public HealthNikhil Behari, Edwin Zhang, Yunfan Zhao, Aparna Taneja et al.NeurIPS 2024 · 39 citations
- f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-DivergencesSiddhant Agarwal, Ishan Durugkar, Peter Stone, Amy ZhangNeurIPS 2023 · 22 citations
- Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AIMalek Mechergui, Sarath SreedharanAAAI 2024 · 18 citations
Builds on1
Related papers
- ORSO: Accelerating Reward Design via Online Reward Selection and Policy OptimizationChen Bo Calvin Zhang, Zhang-Wei Hong, Aldo Pacchiano, Pulkit AgrawalICLR 2025
- Behavior Alignment via Reward Function OptimizationDhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas et al.NeurIPS 2023 · 27 citations
- Programmatic Reward Design by ExampleWeichao Zhou, Wenchao LiAAAI 2022 · 15 citations
- Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RLCharles Packer, Pieter Abbeel, Joseph E. GonzalezNeurIPS 2021 · 22 citations
- Test-driven Reinforcement Learning in Continuous ControlZhao Yu, Xiuping Wu, Liangjun KeAAAI 2026
