Sparse Graphical Memory for Robust Planning
Scott Emmons, Ajay Jain, Michael Laskin, Thanard Kurutach, Pieter Abbeel, Deepak Pathak
摘要
To operate effectively in the real world, agents should be able to act from highdimensional raw sensory input such as images and achieve diverse goals across long time-horizons. Current deep reinforcement and imitation learning methods can learn directly from high-dimensional inputs but do not scale well to long-horizon tasks. In contrast, classical graphical methods like A* search are able to solve long-horizon tasks, but assume that the state space is abstracted away from raw sensory input. Recent works have attempted to combine the strengths of deep learning and classical planning; however, dominant methods in this domain are still quite brittle and scale poorly with the size of the environment. We introduce Sparse Graphical Memory (SGM), a new data structure that stores states and feasible transitions in a sparse memory. SGM aggregates states according to a novel two-way consistency objective, adapting classic state aggregation criteria to goal-conditioned RL: two states are redundant when they are interchangeable both as goals and as starting states. Theoretically, we prove that merging nodes according to two-way consistency leads to an increase in shortest path lengths that scales only linearly with the merging threshold. Experimentally, we show that SGM significantly outperforms current state of the art methods on long horizon, sparsereward visual navigation tasks. Project video and code are available at https: //mishalaskin.github.io/sgm/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language NavigationZhiwei Deng, Karthik Narasimhan, Olga RussakovskyNeurIPS 2020 · 被引用 111 次
- No RL, No Simulation: Learning to Navigate without NavigatingMeera Hahn, Devendra Singh Chaplot, Shubham Tulsiani, Mustafa Mukadam 等NeurIPS 2021 · 被引用 98 次
- World Model as a Graph: Learning Latent Landmarks for PlanningLunjun Zhang, Ge Yang, Bradly C. StadieICML 2021 · 被引用 90 次
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement LearningJunsu Kim, Younggyo Seo, Jinwoo ShinNeurIPS 2021 · 被引用 90 次
- Skill Discovery for Exploration and Planning using Deep Skill GraphsAkhil Bagaria, Jason K. Senthil, George KonidarisICML 2021 · 被引用 73 次
它引用的顶会 Paper2
相关 Paper
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- MemoNav: Working Memory Model for Visual NavigationHongxin Li, Zeyu Wang, Xu Yang, Yuran Yang 等CVPR 2024
- PALMER: Perception - Action Loop with Memory for Long-Horizon PlanningOnur Beker, Mohammad Mohammadi, Amir ZamirNeurIPS 2022 · 被引用 6 次
- Emergence of Spatial Representation in an Actor-Critic Agent with Hippocampus-Inspired Sequence GeneratorXiao-Xiong Lin, Yuk Hoi Yiu, Christian LeiboldICLR 2026 · 被引用 2 次
- Visual Graph Memory with Unsupervised Representation for Visual NavigationObin Kwon, Nuri Kim, Yunho Choi, Hwiyeon Yoo 等ICCV 2021 · 被引用 86 次
