Inferring Implicit Goals Across Differing Task Models
Silvia Tulli, Stylianos Loukas Vasileiou, Mohamed Chetouani, Sarath Sreedharan
Abstract
One of the significant challenges to generating value-aligned behavior is to not only account for the specified user objectives but also any implicit or unspecified user requirements. The existence of such implicit requirements could be particularly common in settings where the user's understanding of the task model may differ from the agent's estimate of the model. Under this scenario, the user may incorrectly expect some agent behavior to be inevitable or guaranteed. This paper addresses such expectation mismatch in the presence of differing models by capturing the possibility of unspecified user subgoal in the context of a task captured as a Markov Decision Process (MDP) and querying for it as required. Our method identifies bottleneck states and uses them as candidates for potential implicit subgoals. We then introduce a querying strategy that will generate the minimal number of queries required to identify a policy guaranteed to achieve the underlying goal. Our empirical evaluations demonstrate the effectiveness of our approach in inferring and achieving unstated goals across various tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8aa5aa55-4e96-4629-9eff-5c03e0b600c4Builds on2
- Quantifying Differences in Reward FunctionsAdam Gleave, Michael Dennis, Shane Legg, Stuart Russell et al.ICLR 2021 · 77 citations
- Hierarchical Expertise-Level Modeling for User Specific Robot-Behavior ExplanationsSarath Sreedharan, Tathagata Chakraborti, Christian Muise, Subbarao KambhampatiAAAI 2020 · 19 citations
Related papers
- Belief-State Query Policies for User-Aligned POMDPsDaniel Bramblett, Siddharth SrivastavaNeurIPS 2024 · 1 citation
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 4 citations
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg et al.ICML 2020 · 81 citations
- Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AIMalek Mechergui, Sarath SreedharanAAAI 2024 · 18 citations
- Explicable Policy SearchZe Gong, Yu ZhangNeurIPS 2022 · 5 citations
