Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AI
Malek Mechergui, Sarath Sreedharan
摘要
While the question of misspecified objectives has gotten much attention in recent years, most works in this area primarily focus on the challenges related to the complexity of the objective specification mechanism (for example, the use of reward functions). However, the complexity of the objective specification mechanism is just one of many reasons why the user may have misspecified their objective. A foundational cause for misspecification that is being overlooked by these works is the inherent asymmetry in human expectations about the agent's behavior and the behavior generated by the agent for the specified objective. To address this, we propose a novel formulation for the objective misspecification problem that builds on the human-aware planning literature, which was originally introduced to support explanation and explicable behavioral generation. Additionally, we propose a first-of-its-kind interactive algorithm that is capable of using information generated under incorrect beliefs about the agent to determine the true underlying goal of the user.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI CompanionsXianzhe Fan, Qing Xiao, Xuhui Zhou, Jiaxin Pei 等CHI 2025 · 被引用 15 次
- Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMsZhaowei Zhang, Fengshuo Bai, Qizhi Chen, Chengdong Ma 等ICLR 2025
- Reducing Goal State Divergence with Environment DesignKelsey Sikes, Sarah Keren, Sarath SreedharanAAAI 2026
它引用的顶会 Paper3
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 被引用 219 次
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum 等AAAI 2023 · 被引用 103 次
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable RepresentationsSarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava 等ICLR 2022 · 被引用 39 次
相关 Paper
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 被引用 4 次
- Hierarchical Expertise-Level Modeling for User Specific Robot-Behavior ExplanationsSarath Sreedharan, Tathagata Chakraborti, Christian Muise, Subbarao KambhampatiAAAI 2020 · 被引用 19 次
- Inferring Implicit Goals Across Differing Task ModelsSilvia Tulli, Stylianos Loukas Vasileiou, Mohamed Chetouani, Sarath SreedharanAAAI 2026
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 被引用 120 次
- Programmatic Reward Design by ExampleWeichao Zhou, Wenchao LiAAAI 2022 · 被引用 15 次
