Goal-Directed Planning via Hindsight Experience Replay
Lorenzo Moro, Amarildo Likmeta, Enrico Prati, Marcello Restelli
Abstract
We consider the problem of goal-directed planning under a deterministic transition model. Monte Carlo Tree Search has shown remarkable performance in solving deterministic control problems. By using function approximators to bias the search of the tree, MCTS has been extended to complex continuous domains, resulting in the AlphaZero family of algorithms. Nonetheless, these algorithms still struggle with control problems with sparse rewards such as goal-directed domains, where a positive reward is awarded only when reaching a goal state. In this work, we extend AlphaZero with Hindsight Experience Replay to tackle complex goal-directed planning tasks. We demonstrate the effectiveness of the proposed approach through an extensive empirical evaluation in several simulated domains, including a novel application to a quantum compiling domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 029dfdb5-17a1-40cd-9488-e9c76d5abc91Cited by top-tier papers4
- Subgoal-based Demonstration Learning for Formal Theorem ProvingXueliang Zhao, Wenda Li, Lingpeng KongICML 2024 · 13 citations
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 7 citations
- Rejecting Hallucinated State Targets during PlanningHarry Zhao, Tristan Sylvain, Romain Laroche, Doina Precup et al.ICML 2025
- SEGO: Sequential Subgoal Optimization for Mathematical Problem-SolvingXueliang Zhao, Xinting Huang, Wei Bi, Lingpeng KongACL 2024
Builds on1
Related papers
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
- Decentralized Monte Carlo Tree Search for Partially Observable Multi-Agent PathfindingAlexey Skrynnik, Anton Andreychuk, Konstantin S. Yakovlev, Aleksandr PanovAAAI 2024 · 21 citations
- Multiagent Gumbel MuZero: Efficient Planning in Combinatorial Action SpacesXiaotian Hao, Jianye Hao, Chenjun Xiao, Kai Li et al.AAAI 2024 · 5 citations
- Convex Regularization in Monte-Carlo Tree SearchTuan Dam, Carlo D'Eramo, Jan Peters, Joni PajarinenICML 2021 · 12 citations
