The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
Gaurav R. Ghosal, Matthew Zurek, Daniel S. Brown, Anca D. Dragan
Abstract
When inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a "rationality coefficient" capturing how much noise or entropy we expect to see in the human behavior. Prior work typically sets the rationality level to a constant value, regardless of the type, or quality, of human feedback. However, in many settings, giving one type of feedback (e.g. a demonstration) may be much more difficult than a different type of feedback (e.g. answering a comparison query). Thus, we expect to see more or less noise depending on the type of human feedback. In this work, we advocate that grounding the rationality coefficient in real data for each feedback type, rather than assuming a default value, has a significant positive effect on reward learning. We test this in both simulated experiments and in a user study with real human feedback. We find that overestimating human rationality can have dire effects on reward learning accuracy and regret. We also find that fitting the rationality coefficient to human data enables better reward learning, even when the human deviates significantly from the noisy-rational choice model due to systematic biases. Further, we find that the rationality level affects the informativeness of each feedback type: surprisingly, demonstrations are not always the most informative---when the human acts very suboptimally, comparisons actually become more informative, even when the rationality level is the same for both. Ultimately, our results emphasize the importance and advantage of paying attention to the assumed human-rationality-level, especially when agents actively learn from multiple types of human feedback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human FeedbackYifu Yuan, Jianye Hao, Yi Ma, Zibin Dong et al.ICLR 2024 · 21 citations
- MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational InferenceRaphaël Baur, Yannick Metz, Maria Gkoulta, Mennatallah El-Assady et al.ICML 2026
- Catching Two Birds with One Stone: Reward Shaping with Dual Random Networks for Balancing Exploration and ExploitationHaozhe Ma, Fangling Li, Jing Yu Lim, Zhengding Luo et al.ICML 2025
- Reward Learning from Multiple Feedback TypesYannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-AssadyICLR 2025
- Reward Learning through Ranking Mean Squared ErrorChaitanya Kharyal, Calarina Muslimani, Matthew TaylorICML 2026
Builds on2
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 113 citations
Related papers
- What Is It You Really Want of Me? Generalized Reward Learning with Biased Beliefs about Domain DynamicsZe Gong, Yu ZhangAAAI 2020 · 14 citations
- On the Sensitivity of Reward Inference to Misspecified Human ModelsJoey Hong, Kush Bhatia, Anca D. DraganICLR 2023
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 7 citations
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths et al.NeurIPS 2022 · 30 citations
- Learning Guarantee of Reward Modeling Using Deep Neural NetworksYuanhang Luo, Yeheng Ge, Ruijian Han, Guohao ShenKDD 2026 · 3 citations
