Learning to Search and Searching to Learn for Generalization in Planning
Michael Aichmüller, Yannik Hesse, Hector Geffner
摘要
Combinatorial generalization remains a central challenge in Deep Reinforcement Learning (DRL). Classical planning provides a simple yet challenging setting to study this problem through explicit relational descriptions, without requiring learning from perception. In sparse-reward domains, standard RL exploration via real-time search is ineffective, and learning-based planning methods often rely on expert demonstrations, hindsight relabeling, or random walks from the goal state. In contrast, planners rely on best-first search methods such as to solve problems from scratch. We propose a self-improving learning framework in combination with a value heuristic represented by a Relational Graph Neural Network: the heuristic guides search, and the resulting search data updates the heuristic via -learning. This loop yields heuristics that can function as general policies and solve new instances even without search, where DRL otherwise fails, as we show on puzzles such as Sokoban, PushWorld, The Witness, and the 2023 International Planning Competition benchmarks. Notably, we demonstrate strong zero-shot generalization: For example, heuristics trained on Blocksworld instances with fewer than blocks successfully solve instances with blocks without search.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu 等NeurIPS 2021 · 被引用 106 次
- Exploration via Elliptical Episodic BonusesMikael Henaff, Roberta Raileanu, Minqi Jiang, Tim RocktäschelNeurIPS 2022 · 被引用 72 次
- Learning General Planning Policies from Small Examples Without SupervisionGuillem Francès, Blai Bonet, Hector GeffnerAAAI 2021 · 被引用 44 次
- Symbolic Network: Generalized Neural Policies for Relational MDPsSankalp Garg, Aniket Bajpai, MausamICML 2020 · 被引用 39 次
相关 Paper
- A Novel Automated Curriculum Strategy to Solve Hard Sokoban Planning InstancesDieqiao Feng, Carla P. Gomes, Bart SelmanNeurIPS 2020 · 被引用 24 次
- Can Q-Learning with Graph Networks Learn a Generalizable Branching Heuristic for a SAT Solver?Vitaly Kurin, Saad Godil, Shimon Whiteson, Bryan CatanzaroNeurIPS 2020 · 被引用 77 次
- Neural Models for Output-Space Invariance in Combinatorial ProblemsYatin Nandwani, Vidit Jain, Mausam, Parag SinglaICLR 2022 · 被引用 3 次
- Graph Neural Network Based Action Ranking for PlanningRajesh Mangannavar, Stefan Lee, Alan Fern, Prasad TadepalliNeurIPS 2025 · 被引用 3 次
- Learning Domain-Independent Heuristics for Grounded and Lifted PlanningDillon Ze Chen, Sylvie Thiébaux, Felipe W. TrevizanAAAI 2024 · 被引用 29 次
