Exploring through Random Curiosity with General Value Functions
Aditya A. Ramesh, Louis Kirsch, Sjoerd van Steenkiste, Jürgen Schmidhuber
摘要
Efficient exploration in reinforcement learning is a challenging problem commonly addressed through intrinsic rewards. Recent prominent approaches are based on state novelty or variants of artificial curiosity. However, directly applying them to partially observable environments can be ineffective and lead to premature dissipation of intrinsic rewards. Here we propose random curiosity with general value functions (RC-GVF), a novel intrinsic reward function that draws upon connections between these distinct approaches. Instead of using only the current observation's novelty or a curiosity bonus for failing to predict precise environment dynamics, RC-GVF derives intrinsic rewards through predicting temporally extended general value functions. We demonstrate that this improves exploration in a hard-exploration diabolical lock problem. Furthermore, RC-GVF significantly outperforms previous methods in the absence of ground-truth episodic counts in the partially observable MiniGrid environments. Panoramic observations on Mini-Grid further boost RC-GVF's performance such that it is competitive to baselines exploiting privileged information in form of episodic counts. * Part of this work was done while the author was a Postdoctoral researcher at IDSIA. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Study of Global and Episodic Bonuses for Exploration in Contextual MDPsMikael Henaff, Minqi Jiang, Roberta RaileanuICML 2023 · 被引用 18 次
- MIMEx: Intrinsic Rewards from Masked Input ModelingToru Lin, Allan JabriNeurIPS 2023 · 被引用 8 次
- Random Latent Exploration for Deep Reinforcement LearningSrinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin 等ICML 2024 · 被引用 8 次
- Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement LearningYiming Fei, Ziming Wang, Rui Yan, Huajin TangICML 2026
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
它引用的顶会 Paper14
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- Kinematic State Abstraction and Provably Efficient Rich-Observation Reinforcement LearningDipendra Misra, Mikael Henaff, Akshay Krishnamurthy, John LangfordICML 2020 · 被引用 158 次
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient LearningAlekh Agarwal, Mikael Henaff, Sham M. Kakade, Wen SunNeurIPS 2020 · 被引用 126 次
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu 等NeurIPS 2021 · 被引用 106 次
相关 Paper
- Go Beyond Imagination: Maximizing Episodic Reachability with World ModelsYao Fu, Run Peng, Honglak LeeICML 2023 · 被引用 1 次
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 被引用 7 次
- Revisiting Intrinsic Reward for Exploration in Procedurally Generated EnvironmentsKaixin Wang, Kuangqi Zhou, Bingyi Kang, Jiashi Feng 等ICLR 2023
- LECO: Learnable Episodic Count for Task-Specific Intrinsic RewardDaeJin Jo, Sungwoong Kim, Daniel Wontae Nam, Taehwan Kwon 等NeurIPS 2022 · 被引用 14 次
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated EnvironmentsDaochen Zha, Wenye Ma, Lei Yuan, Xia Hu 等ICLR 2021 · 被引用 47 次
