Goal Recognition as Reinforcement Learning
Leonardo Amado, Reuth Mirsky, Felipe Meneguzzi
摘要
Most approaches for goal recognition rely on specifications of the possible dynamics of the actor in the environment when pursuing a goal. These specifications suffer from two key issues. First, encoding these dynamics requires careful design by a domain expert, which is often not robust to noise at recognition time. Second, existing approaches often need costly real-time computations to reason about the likelihood of each potential goal. In this paper, we develop a framework that combines model-free reinforcement learning and goal recognition to alleviate the need for careful, manual domain design, and the need for costly online executions. This framework consists of two main stages: Offline learning of policies or utility functions for each potential goal, and online inference. We provide a first instance of this framework using tabular Q-learning for the learning stage, as well as three measures that can be used to perform the inference stage. The resulting instantiation achieves state-of-the-art performance against goal recognizers on standard evaluation domains and superior performance in noisy environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Information Shaping for Enhanced Goal Recognition of Partially-Informed AgentsSarah Keren, Haifeng Xu, Kofi Kwapong, David C. Parkes 等AAAI 2020 · 被引用 14 次
- An LP-Based Approach for Goal Recognition as PlanningLuísa R. de A. Santos, Felipe Meneguzzi, Ramon Fraga Pereira, André Grahl PereiraAAAI 2021 · 被引用 21 次
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing 等NeurIPS 2020 · 被引用 12 次
- Outcome-Driven Reinforcement Learning via Variational InferenceTim G. J. Rudner, Vitchyr Pong, Rowan McAllister, Yarin Gal 等NeurIPS 2021 · 被引用 24 次
- Learning Value Functions from Undirected State-only ExperienceMatthew Chang, Arjun Gupta, Saurabh GuptaICLR 2022 · 被引用 9 次
