Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity
Abhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade, Sergey Levine
Abstract
Reinforcement learning provides an automated framework for learning behaviors from high-level reward specifications, but in practice the choice of reward function can be crucial for good results -- while in principle the reward only needs to specify what the task is, in reality practitioners often need to design more detailed rewards that provide the agent with some hints about how the task should be completed. The idea of this type of ``reward-shaping'' has been often discussed in the literature, and is often a critical part of practical applications, but there is relatively little formal characterization of how the choice of reward shaping can yield benefits in sample complexity. In this work, we build on the framework of novelty-based exploration to provide a simple scheme for incorporating shaped rewards into RL along with an analysis tool to show that particular choices of reward shaping provably improve sample efficiency. We characterize the class of problems where these gains are expected to be significant and show how this can be connected to practical algorithms in the literature. We confirm that these results hold in practice in an experimental evaluation, providing an insight into the mechanisms through which reward shaping can significantly improve the complexity of reinforcement learning while retaining asymptotic performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65199a93-1e5a-470c-91d2-c71ae32a4be9Cited by top-tier papers25
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model FeedbackYufei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian et al.ICML 2024 · 135 citations
- What Makes a Reward Model a Good Teacher? An Optimization PerspectiveNoam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei et al.NeurIPS 2025 · 73 citations
- Reward Shaping for Reinforcement Learning with An Assistant Reward AgentHaozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu et al.ICML 2024 · 34 citations
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
Builds on16
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 308 citations
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen et al.NeurIPS 2020 · 225 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
Related papers
- Exploiting Multiple Abstractions in Episodic RL via Reward ShapingRoberto Cipollone, Giuseppe De Giacomo, Marco Favorito, Luca Iocchi et al.AAAI 2023 · 5 citations
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves et al.AAAI 2023 · 17 citations
- ORSO: Accelerating Reward Design via Online Reward Selection and Policy OptimizationChen Bo Calvin Zhang, Zhang-Wei Hong, Aldo Pacchiano, Pulkit AgrawalICLR 2025
- Exploit Reward Shifting in Value-Based Deep-RL: Optimistic Curiosity-Based Exploration and Conservative Exploitation via Linear Reward ShapingHao Sun, Lei Han, Rui Yang, Xiaoteng Ma et al.NeurIPS 2022 · 46 citations
- Plan-Based Relaxed Reward Shaping for Goal-Directed TasksIngmar Schubert, Ozgur S. Oguz, Marc ToussaintICLR 2021 · 9 citations
