Bootstrapped Reward Shaping
Jacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. Kulkarni
摘要
In reinforcement learning, especially in sparse-reward domains, many environment steps are required to observe reward information. In order to increase the frequency of such observations, "potential-based reward shaping" (PBRS) has been proposed as a method of providing a more dense reward signal while leaving the optimal policy invariant. However, the required "potential function" must be carefully designed with task-dependent knowledge to not deter training performance. In this work, we propose a "bootstrapped" method of reward shaping, termed BSRS, in which the agent's current estimate of the state-value function acts as the potential function for PBRS. We provide convergence proofs for the tabular setting, give insights into training dynamics for deep RL, and show that the proposed method improves training speed in the Atari suite.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 被引用 466 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample ComplexityAbhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade 等NeurIPS 2022 · 被引用 115 次
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 被引用 85 次
- Quantifying Differences in Reward FunctionsAdam Gleave, Michael Dennis, Shane Legg, Stuart Russell 等ICLR 2021 · 被引用 77 次
相关 Paper
- Automatic Reward Shaping from Confounded Offline DataMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2025
- Efficient Potential-based Exploration in Reinforcement Learning using Inverse Dynamic Bisimulation MetricYiming Wang, Ming Yang, Renzhi Dong, Binbin Sun 等NeurIPS 2023 · 被引用 27 次
- Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse RewardsRati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2022 · 被引用 122 次
- Reward Shaping for Reinforcement Learning with An Assistant Reward AgentHaozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu 等ICML 2024 · 被引用 34 次
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves 等AAAI 2023 · 被引用 17 次
