Lune

ICML2026顶会

Learning to Search and Searching to Learn for Generalization in Planning

Michael Aichmüller, Yannik Hesse, Hector Geffner

2026年份

摘要

Combinatorial generalization remains a central challenge in Deep Reinforcement Learning (DRL). Classical planning provides a simple yet challenging setting to study this problem through explicit relational descriptions, without requiring learning from perception. In sparse-reward domains, standard RL exploration via real-time search is ineffective, and learning-based planning methods often rely on expert demonstrations, hindsight relabeling, or random walks from the goal state. In contrast, planners rely on best-first search methods such as A⋆\mathrm{A}^\star to solve problems from scratch. We propose a self-improving WA⋆\mathrm{WA}^\star learning framework in combination with a value heuristic represented by a Relational Graph Neural Network: the heuristic guides search, and the resulting search data updates the heuristic via QQ-learning. This loop yields heuristics that can function as general policies and solve new instances even without search, where DRL otherwise fails, as we show on puzzles such as Sokoban, PushWorld, The Witness, and the 2023 International Planning Competition benchmarks. Notably, we demonstrate strong zero-shot generalization: For example, heuristics trained on Blocksworld instances with fewer than 3030 blocks successfully solve instances with 488488 blocks without search.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext ea70382c-2c1b-4366-b96f-4b5d9bfc21e2

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖