Successor-Predecessor Intrinsic Exploration
Changmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. Gershman
摘要
Exploration is essential in reinforcement learning, particularly in environments where external rewards are sparse. Here we focus on exploration with intrinsic rewards, where the agent transiently augments the external rewards with selfgenerated intrinsic rewards. Although the study of intrinsic rewards has a long history, existing methods focus on composing the intrinsic reward based on measures of future prospects of states, ignoring the information contained in the retrospective structure of transition sequences. Here we argue that the agent can utilise retrospective information to generate explorative behaviour with structure-awareness, facilitating efficient exploration based on global instead of local information. We propose Successor-Predecessor Intrinsic Exploration (SPIE), an exploration algorithm based on a novel intrinsic reward combining prospective and retrospective information. We show that SPIE yields more efficient and ethologically plausible exploratory behaviour in environments with sparse rewards and bottleneck states than competing methods. We also implement SPIE in deep reinforcement learning agents, and show that the resulting agent achieves stronger empirical performance than existing methods on sparse-reward Atari games. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Modelling the control of offline processing with reinforcement learningEleanor Spens, Neil Burgess, Tim E. J. BehrensNeurIPS 2025 · 被引用 2 次
- Hierarchical Successor Representation for Robust TransferChangmin Yu, Máté LengyelICML 2026
它引用的顶会 Paper4
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
- PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement LearningTao Yu, Cuiling Lan, Wenjun Zeng, Mingxiao Feng 等NeurIPS 2021 · 被引用 64 次
- A First-Occupancy Representation for Reinforcement LearningTed Moskovitz, Spencer R. Wilson, Maneesh SahaniICLR 2022 · 被引用 18 次
- Learning State Representations via Retracing in Reinforcement LearningChangmin Yu, Dong Li, Jianye Hao, Jun Wang 等ICLR 2022 · 被引用 9 次
相关 Paper
- Revisiting Intrinsic Reward for Exploration in Procedurally Generated EnvironmentsKaixin Wang, Kuangqi Zhou, Bingyi Kang, Jiashi Feng 等ICLR 2023
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 被引用 48 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 被引用 7 次
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated EnvironmentsDaochen Zha, Wenye Ma, Lei Yuan, Xia Hu 等ICLR 2021 · 被引用 47 次
