Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement Learning
Jacob Adamczyk, Argenis Arriojas, Stas Tiomkin, Rahul V. Kulkarni
Abstract
In reinforcement learning (RL), the ability to utilize prior knowledge from previously solved tasks can allow agents to quickly solve new problems. In some cases, these new problems may be approximately solved by composing the solutions of previously solved primitive tasks (task composition). Otherwise, prior knowledge can be used to adjust the reward function for a new problem, in a way that leaves the optimal policy unchanged but enables quicker learning (reward shaping). In this work, we develop a general framework for reward shaping and task composition in entropy-regularized RL. To do so, we derive an exact relation connecting the optimal soft value functions for two entropy-regularized RL problems with different reward functions and dynamics. We show how the derived relation leads to a general result for reward shaping in entropy-regularized RL. We then generalize this approach to derive an exact relation connecting optimal value functions for the composition of multiple tasks in entropy-regularized RL. We validate these theoretical contributions with experiments showing that reward shaping and task composition lead to faster learning in various settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Bootstrapped Reward ShapingJacob Adamczyk, Volodymyr Makarenko, Stas Tiomkin, Rahul V. KulkarniAAAI 2025 · 7 citations
- Reinforcement Learning for Control of Non-Markovian Cellular Population DynamicsJosiah C. Kratz, Jacob AdamczykICLR 2025
- Highly Efficient Self-Adaptive Reward Shaping for Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima et al.ICLR 2025
Builds on4
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 72 citations
- A Boolean Task Algebra for Reinforcement LearningGeraud Nangue Tasse, Steven James, Benjamin RosmanNeurIPS 2020 · 71 citations
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 23 citations
Related papers
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah et al.AAAI 2021 · 54 citations
- Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum LearningJin Hwa Lee, Stefano Sarao Mannelli, Andrew M. SaxeICML 2024 · 15 citations
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves et al.AAAI 2023 · 17 citations
- Learning Task-Distribution Reward Shaping with Meta-LearningHaosheng Zou, Tongzheng Ren, Dong Yan, Hang Su et al.AAAI 2021 · 19 citations
- Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample ComplexityAbhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade et al.NeurIPS 2022 · 115 citations
