Learning to Search and Searching to Learn for Generalization in Planning
Michael Aichmüller, Yannik Hesse, Hector Geffner
Abstract
Combinatorial generalization remains a central challenge in Deep Reinforcement Learning (DRL). Classical planning provides a simple yet challenging setting to study this problem through explicit relational descriptions, without requiring learning from perception. In sparse-reward domains, standard RL exploration via real-time search is ineffective, and learning-based planning methods often rely on expert demonstrations, hindsight relabeling, or random walks from the goal state. In contrast, planners rely on best-first search methods such as to solve problems from scratch. We propose a self-improving learning framework in combination with a value heuristic represented by a Relational Graph Neural Network: the heuristic guides search, and the resulting search data updates the heuristic via -learning. This loop yields heuristics that can function as general policies and solve new instances even without search, where DRL otherwise fails, as we show on puzzles such as Sokoban, PushWorld, The Witness, and the 2023 International Planning Competition benchmarks. Notably, we demonstrate strong zero-shot generalization: For example, heuristics trained on Blocksworld instances with fewer than blocks successfully solve instances with blocks without search.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea70382c-2c1b-4366-b96f-4b5d9bfc21e2Builds on12
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu et al.NeurIPS 2021 · 106 citations
- Exploration via Elliptical Episodic BonusesMikael Henaff, Roberta Raileanu, Minqi Jiang, Tim RocktäschelNeurIPS 2022 · 72 citations
- Learning General Planning Policies from Small Examples Without SupervisionGuillem Francès, Blai Bonet, Hector GeffnerAAAI 2021 · 44 citations
- Symbolic Network: Generalized Neural Policies for Relational MDPsSankalp Garg, Aniket Bajpai, MausamICML 2020 · 39 citations
Related papers
- A Novel Automated Curriculum Strategy to Solve Hard Sokoban Planning InstancesDieqiao Feng, Carla P. Gomes, Bart SelmanNeurIPS 2020 · 24 citations
- Can Q-Learning with Graph Networks Learn a Generalizable Branching Heuristic for a SAT Solver?Vitaly Kurin, Saad Godil, Shimon Whiteson, Bryan CatanzaroNeurIPS 2020 · 77 citations
- Neural Models for Output-Space Invariance in Combinatorial ProblemsYatin Nandwani, Vidit Jain, Mausam, Parag SinglaICLR 2022 · 3 citations
- Graph Neural Network Based Action Ranking for PlanningRajesh Mangannavar, Stefan Lee, Alan Fern, Prasad TadepalliNeurIPS 2025 · 3 citations
- Learning Domain-Independent Heuristics for Grounded and Lifted PlanningDillon Ze Chen, Sylvie Thiébaux, Felipe W. TrevizanAAAI 2024 · 29 citations
