Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch
Malek Mechergui, Sarath Sreedharan
Abstract
Detecting and handling misspecified objectives, such as reward functions, has been widely recognized as one of the central challenges within the domain of Artificial Intelligence (AI) safety research. However, even with the recognition of the importance of this problem, we are unaware of any works that attempt to provide a clear definition for what constitutes (a) misspecified objectives and (b) successfully resolving such misspecifications. In this work, we use the theory of mind, i.e., the human user's beliefs about the AI agent, as a basis to develop a formal explanatory framework called Expectation Alignment (EAL) to understand the objective misspecification and its causes. Our EAL framework not only acts as an explanatory framework for existing works but also provides us with concrete insights into the limitations of existing methods to handle reward misspecification and novel solution strategies. We use these insights to propose a new interactive algorithm that uses the specified reward to infer potential user expectations about the system behavior. We show how one can efficiently implement this algorithm by mapping the inference problem into linear programs. We evaluate our method on a set of standard Markov Decision Process (MDP) benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3de88251-be87-4920-bdb1-5d8ef48b9eb6Builds on6
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho et al.NeurIPS 2021 · 107 citations
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum et al.AAAI 2023 · 103 citations
- What Is It You Really Want of Me? Generalized Reward Learning with Biased Beliefs about Domain DynamicsZe Gong, Yu ZhangAAAI 2020 · 14 citations
Related papers
- Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AIMalek Mechergui, Sarath SreedharanAAAI 2024 · 18 citations
- Discovering Implicit Large Language Model Alignment ObjectivesEdward Chen, Sanmi Koyejo, Carlos GuestrinICML 2026
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 2 citations
- Addressing and Visualizing Misalignments in Human Task-Solving TrajectoriesSejin Kim, Hosung Lee, Sundong KimKDD 2025 · 2 citations
- On the Sensitivity of Reward Inference to Misspecified Human ModelsJoey Hong, Kush Bhatia, Anca D. DraganICLR 2023
