A Novel Automated Curriculum Strategy to Solve Hard Sokoban Planning Instances
Dieqiao Feng, Carla P. Gomes, Bart Selman
Abstract
In recent years, we have witnessed tremendous progress in deep reinforcement learning (RL) for tasks such as Go, Chess, video games, and robot control. Nevertheless, other combinatorial domains, such as AI planning, still pose considerable challenges for RL approaches. The key difficulty in those domains is that a positive reward signal becomes exponentially rare as the minimal solution length increases. So, an RL approach loses its training signal. There has been promising recent progress by using a curriculum-driven learning approach that is designed to solve a single hard instance. We present a novel automated curriculum approach that dynamically selects from a pool of unlabeled training instances of varying task complexity guided by our difficulty quantum momentum strategy. We show how the smoothness of the task hardness impacts the final learning results. In particular, as the size of the instance pool increases, the ``hardness gap'' decreases, which facilitates a smoother automated curriculum based learning process. Our automated curriculum approach dramatically improves upon the previous approaches. We show our results on Sokoban, which is a traditional PSPACE-complete planning problem and presents a great challenge even for specialized solvers. Our RL agent can solve hard instances that are far out of reach for any previous state-of-the-art Sokoban solver. In particular, our approach can uncover plans that require hundreds of steps, while the best previous search methods would take many years of computing time to solve such instances. In addition, we show that we can further boost the RL performance with an intricate coupling of our automated curriculum approach with a curiosity-driven search strategy and a graph neural net representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e2d012f-6556-4e27-8d96-9fece3ced9b0Cited by top-tier papers5
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 123 citations
- Environment Generation for Zero-Shot Compositional Reinforcement LearningIzzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi et al.NeurIPS 2021 · 50 citations
- SeeA*: Efficient Exploration-Enhanced A* Search by Selective SamplingDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2024 · 4 citations
- Left Heavy Tails and the Effectiveness of the Policy and Value Networks in DNN-based best-first search for Sokoban PlanningDieqiao Feng, Carla P. Gomes, Bart SelmanNeurIPS 2022 · 3 citations
- Weighted Sampling without Replacement for Deep Top-k ClassificationDieqiao Feng, Yuanqi Du, Carla P. Gomes, Bart SelmanICML 2023
Related papers
- Learning to Search and Searching to Learn for Generalization in PlanningMichael Aichmüller, Yannik Hesse, Hector GeffnerICML 2026
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 83 citations
- CQM: Curriculum Reinforcement Learning with a Quantized World ModelSeungjae Lee, Daesol Cho, Jonghae Park, H. Jin KimNeurIPS 2023 · 18 citations
- Curriculum Reinforcement Learning via Constrained Optimal TransportPascal Klink, Haoyi Yang, Carlo D'Eramo, Jan Peters et al.ICML 2022 · 44 citations
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
