Successor Feature Landmarks for Long-Horizon Goal-Conditioned Reinforcement Learning
Christopher Hoang, Sungryull Sohn, Jongwook Choi, Wilka Carvalho, Honglak Lee
Abstract
Operating in the real-world often requires agents to learn about a complex environment and apply this understanding to achieve a breadth of goals. This problem, known as goal-conditioned reinforcement learning (GCRL), becomes especially challenging for long-horizon goals. Current methods have tackled this problem by augmenting goal-conditioned policies with graph-based planning algorithms. However, they struggle to scale to large, high-dimensional state spaces and assume access to exploration mechanisms for efficiently collecting training data. In this work, we introduce Successor Feature Landmarks (SFL), a framework for exploring large, high-dimensional environments so as to obtain a policy that is proficient for any goal. SFL leverages the ability of successor features (SF) to capture transition dynamics, using it to drive exploration by estimating state-novelty and to enable high-level planning by abstracting the state-space as a non-parametric landmarkbased graph. We further exploit SF to directly compute a goal-conditioned policy for inter-landmark traversal, which we use to execute plans to "frontier" landmarks at the edge of the explored state space. We show in our experiments on MiniGrid and ViZDoom that SFL enables efficient exploration of large, high-dimensional state spaces and outperforms state-of-the-art baselines on long-horizon GCRL tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1f3fa0d-e96d-4933-9f11-82f21b46b2a1Cited by top-tier papers12
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach et al.NeurIPS 2025 · 60 citations
- CQM: Curriculum Reinforcement Learning with a Quantized World ModelSeungjae Lee, Daesol Cho, Jonghae Park, H. Jin KimNeurIPS 2023 · 18 citations
- Unsupervised Object Interaction Learning with Counterfactual Dynamics ModelsJongwook Choi, Sungtae Lee, Xinyu Wang, Sungryull Sohn et al.AAAI 2024 · 6 citations
- Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement LearningSeungyul Han, Jaebak Hwang, Sanghyeon Lee, Jeongmo KimICLR 2026 · 3 citations
Builds on5
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- Sparse Graphical Memory for Robust PlanningScott Emmons, Ajay Jain, Michael Laskin, Thanard Kurutach et al.NeurIPS 2020 · 60 citations
- Hallucinative Topological Memory for Zero-Shot Visual PlanningKara Liu, Thanard Kurutach, Christine Tung, Pieter Abbeel et al.ICML 2020 · 49 citations
- Neural Topological SLAM for Visual NavigationDevendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, Saurabh GuptaCVPR 2020
Related papers
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement LearningJunsu Kim, Younggyo Seo, Jinwoo ShinNeurIPS 2021 · 90 citations
- World Model as a Graph: Learning Latent Landmarks for PlanningLunjun Zhang, Ge Yang, Bradly C. StadieICML 2021 · 90 citations
- C-Planning: An Automatic Curriculum for Learning Goal-Reaching TasksTianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine et al.ICLR 2022 · 19 citations
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu et al.ICLR 2022 · 28 citations
- Flattening Hierarchies with Policy BootstrappingJohn L. Zhou, Jonathan C. KaoNeurIPS 2025 · 8 citations
