Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch
Malek Mechergui, Sarath Sreedharan
摘要
Detecting and handling misspecified objectives, such as reward functions, has been widely recognized as one of the central challenges within the domain of Artificial Intelligence (AI) safety research. However, even with the recognition of the importance of this problem, we are unaware of any works that attempt to provide a clear definition for what constitutes (a) misspecified objectives and (b) successfully resolving such misspecifications. In this work, we use the theory of mind, i.e., the human user's beliefs about the AI agent, as a basis to develop a formal explanatory framework called Expectation Alignment (EAL) to understand the objective misspecification and its causes. Our EAL framework not only acts as an explanatory framework for existing works but also provides us with concrete insights into the limitations of existing methods to handle reward misspecification and novel solution strategies. We use these insights to propose a new interactive algorithm that uses the specified reward to infer potential user expectations about the system behavior. We show how one can efficiently implement this algorithm by mapping the inference problem into linear programs. We evaluate our method on a set of standard Markov Decision Process (MDP) benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 被引用 293 次
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 被引用 219 次
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum 等AAAI 2023 · 被引用 103 次
- What Is It You Really Want of Me? Generalized Reward Learning with Biased Beliefs about Domain DynamicsZe Gong, Yu ZhangAAAI 2020 · 被引用 14 次
相关 Paper
- Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AIMalek Mechergui, Sarath SreedharanAAAI 2024 · 被引用 18 次
- Discovering Implicit Large Language Model Alignment ObjectivesEdward Chen, Sanmi Koyejo, Carlos GuestrinICML 2026
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 被引用 2 次
- Addressing and Visualizing Misalignments in Human Task-Solving TrajectoriesSejin Kim, Hosung Lee, Sundong KimKDD 2025 · 被引用 2 次
- On the Sensitivity of Reward Inference to Misspecified Human ModelsJoey Hong, Kush Bhatia, Anca D. DraganICLR 2023
