Lune

NeurIPS2022顶会

Learning Generalized Policy Automata for Relational Stochastic Shortest Path Problems

Rushang Karia, Rashmeet Kaur Nayyar, Siddharth Srivastava

2022年份
3被引次数
1顶会引用

摘要

Several goal-oriented problems in the real-world can be naturally expressed as Stochastic Shortest Path Problems (SSPs). However, the computational complexity of solving SSPs makes finding solutions to even moderately sized problems intractable. Currently, existing state-of-the-art planners and heuristics often fail to exploit knowledge learned from solving other instances. This paper presents an approach for learning Generalized Policy Automata (GPA): non-deterministic partial policies that can be used to catalyze the solution process. GPAs are learned using relational, feature-based abstractions, which makes them applicable on broad classes of related problems with different object names and quantities. Theoretical analysis of this approach shows that it guarantees completeness and hierarchical optimality. Empirical analysis shows that this approach effectively learns broadly applicable policy knowledge in a few-shot fashion and significantly outperforms state-of-the-art SSP solvers on test problems whose object counts are far greater than those used during training. Running example: The planetary rover example can be expressed using a domain that consists of parameterized predicates connected(l x , l y ), in-rover(r x ), rock-at(r x , l x ), and actions load(r x , l x ), unload(r x , l x ), and move(l x , l y ). Object types can be denoted using unary predicates location(l x ) and rock(r x ). l x , l y , and r x are parameters that can be instantiated with different locations and

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper3

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖