Learning Safe Numeric Action Models
Argaman Mordoch, Brendan Juba, Roni Stern
摘要
Powerful domain-independent planners have been developed to solve various types of planning problems. These planners often require a model of the acting agent's actions, given in some planning domain description language. Yet obtaining such an action model is a notoriously hard task. This task is even more challenging in mission-critical domains, where a trial-and-error approach to learning how to act is not an option. In such domains, the action model used to generate plans must be safe, in the sense that plans generated with it must be applicable and achieve their goals. Learning safe action models for planning has been recently explored for domains in which states are sufficiently described with Boolean variables. In this work, we go beyond this limitation and propose the Numeric Safe Action Model Learning (N-SAM) algorithm. N-SAM runs in time that is polynomial in the number of observations and, under certain conditions, is guaranteed to return safe action models. We analyze its worst-case sample complexity, which may be intractable for some domains. Empirically, however, N-SAM can quickly learn a safe action model that can solve most problems in the domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Learning Probably Approximately Complete and Safe Action Models for Stochastic WorldsBrendan Juba, Roni SternAAAI 2022 · 被引用 18 次
相关 Paper
- Learning Safe Action Models with Partial ObservabilityHai S. Le, Brendan Juba, Roni SternAAAI 2024 · 被引用 7 次
- Learning Planning Domains from Non-redundant Fully-Observed Traces: Theoretical Foundations and Complexity AnalysisPascal Bachor, Gregor BehnkeAAAI 2024 · 被引用 5 次
- Learning with Safety Constraints: Sample Complexity of Reinforcement Learning for Constrained MDPsAria HasanzadeZonuzy, Archana Bura, Dileep M. Kalathil, Srinivas ShakkottaiAAAI 2021 · 被引用 46 次
- NaRuto: Automatically Acquiring Planning Models from Narrative TextsRuiqi Li, Leyang Cui, Songtuan Lin, Patrik HaslumAAAI 2024 · 被引用 8 次
- Learning General Planning Policies from Small Examples Without SupervisionGuillem Francès, Blai Bonet, Hector GeffnerAAAI 2021 · 被引用 44 次
