Revisiting Intrinsic Reward for Exploration in Procedurally Generated Environments
Kaixin Wang, Kuangqi Zhou, Bingyi Kang, Jiashi Feng, Shuicheng Yan
Abstract
Exploration under sparse rewards remains a key challenge in deep reinforcement learning. Recently, studying exploration in procedurally-generated environments has drawn increasing attention. Existing works generally combine lifelong intrinsic rewards and episodic intrinsic rewards to encourage exploration. Though various lifelong and episodic intrinsic rewards have been proposed, the individual contributions of the two kinds of intrinsic rewards to improving exploration are barely investigated. To bridge this gap, we disentangle these two parts and conduct ablative experiments. We consider lifelong and episodic intrinsic rewards used in prior works, and compare the performance of all lifelong-episodic combinations on the commonly used MiniGrid benchmark. Experimental results show that only using episodic intrinsic rewards can match or surpass prior state-of-the-art methods. On the other hand, only using lifelong intrinsic rewards hardly makes progress in exploration. This demonstrates that episodic intrinsic reward is more crucial than lifelong one in boosting exploration. Moreover, we find through experimental analysis that the lifelong intrinsic reward does not accurately reflect the novelty of states, which explains why it does not help much in improving exploration.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f8207fa1-594d-4c55-90fd-8f5bddae2023Cited by top-tier papers3
- A Study of Global and Episodic Bonuses for Exploration in Contextual MDPsMikael Henaff, Minqi Jiang, Roberta RaileanuICML 2023 · 18 citations
- Improving Intrinsic Exploration by Creating Stationary ObjectivesRoger Creus Castanyer, Joshua Romoff, Glen BersethICLR 2024 · 4 citations
- Abstract and Explore: A Novel Behavioral Metric with Cyclic Dynamics in Reinforcement LearningAnjie Zhu, Peng-Fei Zhang, Ruihong Qiu, Zetao Zheng et al.AAAI 2024 · 1 citation
Related papers
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated EnvironmentsDaochen Zha, Wenye Ma, Lei Yuan, Xia Hu et al.ICLR 2021 · 47 citations
- Go Beyond Imagination: Maximizing Episodic Reachability with World ModelsYao Fu, Run Peng, Honglak LeeICML 2023 · 1 citation
- Successor-Predecessor Intrinsic ExplorationChangmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. GershmanNeurIPS 2023 · 12 citations
- Improving Intrinsic Exploration with Language AbstractionsJesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang et al.NeurIPS 2022 · 81 citations
