Learning General Planning Policies from Small Examples Without Supervision
Guillem Francès, Blai Bonet, Hector Geffner
摘要
Generalized planning is concerned with the computation of general policies that solve multiple instances of a planning domain all at once. It has been recently shown that these policies can be computed in two steps: first, a suitable abstraction in the form of a qualitative numerical planning problem (QNP) is learned from sample plans, then the general policies are obtained from the learned QNP using a planner. In this work, we introduce an alternative approach for computing more expressive general policies which does not require sample plans or a QNP planner. The new formulation is very simple and can be cast in terms that are more standard in machine learning: a large but finite pool of features is defined from the predicates in the planning examples using a general grammar, and a small subset of features is sought for separating “good” from “bad” state transitions, and goals from non-goals. The problems of finding such a “separating surface” while labeling the transitions as “good” or “bad” are jointly addressed as a single combinatorial optimization problem expressed as a Weighted Max-SAT problem. The advantage of looking for the simplest policy in the given feature space that solves the given examples, possibly non-optimally, is that many domains have no general, compact policies that are optimal. The approach yields general policies for a number of benchmark domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- An Automatic Sound and Complete Abstraction Method for Generalized Planning with Baggable TypesHao Dong, Zheyuan Shi, Hemeng Zeng, Yongmei LiuAAAI 2025 · 被引用 3 次
- Learning Generalized Policy Automata for Relational Stochastic Shortest Path ProblemsRushang Karia, Rashmeet Kaur Nayyar, Siddharth SrivastavaNeurIPS 2022 · 被引用 3 次
- Satisficing and Optimal Generalised Planning via Goal RegressionDillon Z. Chen, Till Hofmann, Toryn Q. Klassen, Sheila A. McIlraithAAAI 2026 · 被引用 1 次
- First-Order Representation Languages for Goal-Conditioned RLSimon Ståhlberg, Hector GeffnerAAAI 2026 · 被引用 1 次
- Learning More Expressive General Policies for Classical Planning DomainsSimon Ståhlberg, Blai Bonet, Hector GeffnerAAAI 2025
它引用的顶会 Paper3
- Few-Shot Bayesian Imitation Learning with Logical Program PoliciesTom Silver, Kelsey R. Allen, Alex K. Lew, Leslie Pack Kaelbling 等AAAI 2020 · 被引用 57 次
- Symbolic Network: Generalized Neural Policies for Relational MDPsSankalp Garg, Aniket Bajpai, MausamICML 2020 · 被引用 39 次
- General Policies, Representations, and Planning WidthBlai Bonet, Hector GeffnerAAAI 2021 · 被引用 28 次
相关 Paper
- Discovering State and Action Abstractions for Generalized Task and Motion PlanningAidan Curtis, Tom Silver, Joshua B. Tenenbaum, Tomás Lozano-Pérez 等AAAI 2022 · 被引用 36 次
- Predicate Invention for Bilevel PlanningTom Silver, Rohan Chitnis, Nishanth Kumar, Willie McClinton 等AAAI 2023 · 被引用 73 次
- Generalized Planning with Positive and Negative ExamplesJavier Segovia-Aguas, Sergio Jiménez, Anders JonssonAAAI 2020 · 被引用 11 次
- Learning Generalized Relational Heuristic Networks for Model-Agnostic PlanningRushang Karia, Siddharth SrivastavaAAAI 2021 · 被引用 49 次
- Learning Safe Numeric Action ModelsArgaman Mordoch, Brendan Juba, Roni SternAAAI 2023 · 被引用 6 次
