Successor-Predecessor Intrinsic Exploration
Changmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. Gershman
Abstract
Exploration is essential in reinforcement learning, particularly in environments where external rewards are sparse. Here we focus on exploration with intrinsic rewards, where the agent transiently augments the external rewards with selfgenerated intrinsic rewards. Although the study of intrinsic rewards has a long history, existing methods focus on composing the intrinsic reward based on measures of future prospects of states, ignoring the information contained in the retrospective structure of transition sequences. Here we argue that the agent can utilise retrospective information to generate explorative behaviour with structure-awareness, facilitating efficient exploration based on global instead of local information. We propose Successor-Predecessor Intrinsic Exploration (SPIE), an exploration algorithm based on a novel intrinsic reward combining prospective and retrospective information. We show that SPIE yields more efficient and ethologically plausible exploratory behaviour in environments with sparse rewards and bottleneck states than competing methods. We also implement SPIE in deep reinforcement learning agents, and show that the resulting agent achieves stronger empirical performance than existing methods on sparse-reward Atari games. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0592f34c-e907-424b-bd5f-ffbfbf545de7Cited by top-tier papers2
- Modelling the control of offline processing with reinforcement learningEleanor Spens, Neil Burgess, Tim E. J. BehrensNeurIPS 2025 · 2 citations
- Hierarchical Successor Representation for Robust TransferChangmin Yu, Máté LengyelICML 2026
Builds on4
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement LearningTao Yu, Cuiling Lan, Wenjun Zeng, Mingxiao Feng et al.NeurIPS 2021 · 64 citations
- A First-Occupancy Representation for Reinforcement LearningTed Moskovitz, Spencer R. Wilson, Maneesh SahaniICLR 2022 · 18 citations
- Learning State Representations via Retracing in Reinforcement LearningChangmin Yu, Dong Li, Jianye Hao, Jun Wang et al.ICLR 2022 · 9 citations
Related papers
- Revisiting Intrinsic Reward for Exploration in Procedurally Generated EnvironmentsKaixin Wang, Kuangqi Zhou, Bingyi Kang, Jiashi Feng et al.ICLR 2023
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 48 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 7 citations
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated EnvironmentsDaochen Zha, Wenye Ma, Lei Yuan, Xia Hu et al.ICLR 2021 · 47 citations
