World Model as a Graph: Learning Latent Landmarks for Planning
Lunjun Zhang, Ge Yang, Bradly C. Stadie
Abstract
Planning, the ability to analyze the structure of a problem in the large and decompose it into interrelated subproblems, is a hallmark of human intelligence. While deep reinforcement learning (RL) has shown great promise for solving relatively straightforward control tasks, it remains an open problem how to best incorporate planning into existing deep RL paradigms to handle increasingly complex environments. One prominent framework, Model-Based RL, learns a world model and plans using step-by-step virtual rollouts. This type of world model quickly diverges from reality when the planning horizon increases, thus struggling at long-horizon planning. How can we learn world models that endow agents with the ability to do temporally extended reasoning? In this work, we propose to learn graph-structured world models composed of sparse, multi-step transitions. We devise a novel algorithm to learn latent landmarks that are scattered (in terms of reachability) across the goal space as the nodes on the graph. In this same graph, the edges are the reachability estimates distilled from Q-functions. On a variety of high-dimensional continuous control tasks ranging from robotic manipulation to navigation, we demonstrate that our method, named L 3 P , significantly outperforms prior work, and is oftentimes the only method capable of leveraging both the robustness of model-free RL and generalization of graph-search algorithms. We believe our work is an important step towards scalable planning in reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cf8069e-bed4-4753-8744-976660cba5a2Cited by top-tier papers29
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement LearningJunsu Kim, Younggyo Seo, Jinwoo ShinNeurIPS 2021 · 90 citations
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 44 citations
- Subgoal Search For Complex Reasoning TasksKonrad Czechowski, Tomasz Odrzygózdz, Marek Zbysinski, Michal Zawalski et al.NeurIPS 2021 · 41 citations
- DHRL: A Graph-Based Approach for Long-Horizon and Sparse Hierarchical Reinforcement LearningSeungjae Lee, Jigang Kim, Inkyu Jang, H. Jin KimNeurIPS 2022 · 33 citations
Builds on8
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie et al.ICML 2020 · 145 citations
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou et al.NeurIPS 2020 · 96 citations
Related papers
- Skill Discovery for Exploration and Planning using Deep Skill GraphsAkhil Bagaria, Jason K. Senthil, George KonidarisICML 2021 · 73 citations
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Flexible and Efficient Long-Range Planning Through Curious ExplorationAidan Curtis, Minjian Xin, Dilip Arumugam, Kevin T. Feigelis et al.ICML 2020 · 7 citations
- Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement LearningHarry Zhao, Safa Alver, Harm van Seijen, Romain Laroche et al.ICLR 2024 · 5 citations
- Successor Feature Landmarks for Long-Horizon Goal-Conditioned Reinforcement LearningChristopher Hoang, Sungryull Sohn, Jongwook Choi, Wilka Carvalho et al.NeurIPS 2021 · 47 citations
