Redeeming intrinsic rewards via constrained optimization
Eric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit Agrawal
摘要
State-of-the-art reinforcement learning (RL) algorithms typically use random sampling (e.g., -greedy) for exploration, but this method fails on hard exploration tasks like Montezuma's Revenge. To address the challenge of exploration, prior works incentivize exploration by rewarding the agent when it visits novel states. Such intrinsic rewards (also called exploration bonus or curiosity) often lead to excellent performance on hard exploration tasks. However, on easy exploration tasks, the agent gets distracted by intrinsic rewards and performs unnecessary exploration even when sufficient task (also called extrinsic) reward is available. Consequently, such an overly curious agent performs worse than an agent trained with only task reward. Such inconsistency in performance across tasks prevents the widespread use of intrinsic rewards with RL algorithms. We propose a principled constrained optimization procedure called Extrinsic-Intrinsic Policy Optimization (EIPO) that automatically tunes the importance of the intrinsic reward: it suppresses the intrinsic reward when exploration is unnecessary and increases it when exploration is required. The results is superior exploration that does not require manual tuning in balancing the intrinsic reward against the task reward. Consistent performance gains across sixty-one ATARI games validate our claim. The code is available at https://github.com/Improbable-AI/eipo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
- Random Latent Exploration for Deep Reinforcement LearningSrinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin 等ICML 2024 · 被引用 8 次
- Breadcrumbs to the Goal: Supervised Goal Selection from Human-in-the-Loop FeedbackMarcel Torne Villasevil, Max Balsells, Zihan Wang, Samedh Desai 等NeurIPS 2023 · 被引用 3 次
- Off-policy Reinforcement Learning with Model-based Exploration AugmentationLikun Wang, Xiangteng Zhang, Yinuo Wang, Guojian Zhan 等NeurIPS 2025 · 被引用 3 次
- Action-Dependent Optimality-Preserving Reward ShapingGrant C. Forbes, Jianxun Wang, Leonardo Villalobos-Arias, Arnav Jhala 等ICML 2025
它引用的顶会 Paper4
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 被引用 241 次
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu 等ICML 2020 · 被引用 87 次
相关 Paper
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel 等ICLR 2025
- Continuously Discovering Novel Strategies via Reward-Switching Policy OptimizationZihan Zhou, Wei Fu, Bingliang Zhang, Yi WuICLR 2022 · 被引用 34 次
- Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICML 2023 · 被引用 17 次
- Successor-Predecessor Intrinsic ExplorationChangmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. GershmanNeurIPS 2023 · 被引用 12 次
- Going Beyond Heuristics by Imposing Policy Improvement as a ConstraintChi-Chang Lee, Zhang-Wei Hong, Pulkit AgrawalNeurIPS 2024 · 被引用 2 次
