Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AI
Malek Mechergui, Sarath Sreedharan
Abstract
While the question of misspecified objectives has gotten much attention in recent years, most works in this area primarily focus on the challenges related to the complexity of the objective specification mechanism (for example, the use of reward functions). However, the complexity of the objective specification mechanism is just one of many reasons why the user may have misspecified their objective. A foundational cause for misspecification that is being overlooked by these works is the inherent asymmetry in human expectations about the agent's behavior and the behavior generated by the agent for the specified objective. To address this, we propose a novel formulation for the objective misspecification problem that builds on the human-aware planning literature, which was originally introduced to support explanation and explicable behavioral generation. Additionally, we propose a first-of-its-kind interactive algorithm that is capable of using information generated under incorrect beliefs about the agent to determine the true underlying goal of the user.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 356a0ad9-c484-4a3c-b5ac-590ea52b6a52Cited by top-tier papers3
- User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI CompanionsXianzhe Fan, Qing Xiao, Xuhui Zhou, Jiaxin Pei et al.CHI 2025 · 15 citations
- Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMsZhaowei Zhang, Fengshuo Bai, Qizhi Chen, Chengdong Ma et al.ICLR 2025
- Reducing Goal State Divergence with Environment DesignKelsey Sikes, Sarah Keren, Sarath SreedharanAAAI 2026
Builds on3
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum et al.AAAI 2023 · 103 citations
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable RepresentationsSarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava et al.ICLR 2022 · 39 citations
Related papers
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 4 citations
- Hierarchical Expertise-Level Modeling for User Specific Robot-Behavior ExplanationsSarath Sreedharan, Tathagata Chakraborti, Christian Muise, Subbarao KambhampatiAAAI 2020 · 19 citations
- Inferring Implicit Goals Across Differing Task ModelsSilvia Tulli, Stylianos Loukas Vasileiou, Mohamed Chetouani, Sarath SreedharanAAAI 2026
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- Programmatic Reward Design by ExampleWeichao Zhou, Wenchao LiAAAI 2022 · 15 citations
