Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse Rewards
Rati Devidze, Parameswaran Kamalaruban, Adish Singla
Abstract
We study the problem of reward shaping to accelerate the training process of a reinforcement learning agent. Existing works have considered a number of different reward shaping formulations; however, they either require external domain knowledge or fail in environments with extremely sparse rewards. In this paper, we propose a novel framework, Exploration-Guided Reward Shaping (E XPLO RS), that operates in a fully self-supervised manner and can accelerate an agent’s learning even in sparse-reward environments. The key idea of E XPLO RS is to learn an intrinsic reward function in combination with exploration-based bonuses to maximize the agent’s utility w.r.t. extrinsic rewards. We theoretically showcase the usefulness of our reward shaping framework in a special family of MDPs. Experimental results on several environments with sparse/noisy reward signals demonstrate the effectiveness of E XPLO RS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e37204f-90f7-4854-a93d-d234658f109fCited by top-tier papers21
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage ShapingThanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai et al.ICLR 2026 · 55 citations
- Reward Shaping for Reinforcement Learning with An Assistant Reward AgentHaozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu et al.ICML 2024 · 34 citations
- Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner CollaborationBowei He, Minda Hu, Zenan Xu, Hongru WANG et al.ICML 2026 · 9 citations
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima et al.NeurIPS 2025 · 9 citations
- GUIDE: Real-Time Human-Shaped AgentsLingyu Zhang, Zhengran Ji, Nicholas R. Waytowich, Boyuan ChenNeurIPS 2024 · 9 citations
Builds on3
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 87 citations
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 60 citations
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah et al.AAAI 2021 · 54 citations
Related papers
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves et al.AAAI 2023 · 17 citations
- Action-Dependent Optimality-Preserving Reward ShapingGrant C. Forbes, Jianxun Wang, Leonardo Villalobos-Arias, Arnav Jhala et al.ICML 2025
- Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICML 2023 · 17 citations
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 7 citations
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 48 citations
