Learning Guidance Rewards with Trajectory-space Smoothing
Tanmay Gangwani, Yuan Zhou, Jian Peng
Abstract
Long-term temporal credit assignment is an important challenge in deep reinforcement learning (RL). It refers to the ability of the agent to attribute actions to consequences that may occur after a long time interval. Existing policy-gradient and Q-learning algorithms typically rely on dense environmental rewards that provide rich short-term supervision and help with credit assignment. However, they struggle to solve tasks with delays between an action and the corresponding rewarding feedback. To make credit assignment easier, recent works have proposed algorithms to learn dense "guidance" rewards that could be used in place of the sparse or delayed environmental rewards. This paper is in the same vein -- starting with a surrogate RL objective that involves smoothing in the trajectory-space, we arrive at a new algorithm for learning guidance rewards. We show that the guidance rewards have an intuitive interpretation, and can be obtained without training any additional neural networks. Due to the ease of integration, we use the guidance rewards in a few popular algorithms (Q-learning, Actor-Critic, Distributional-RL) and present results in single-agent and multi-agent tasks that elucidate the benefit of our approach when the environmental rewards are sparse or delayed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 831e265a-7e7a-4cbd-a635-2d6137826833Cited by top-tier papers12
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 74 citations
- Off-Policy Reinforcement Learning with Delayed RewardsBeining Han, Zhizhou Ren, Zuofan Wu, Yuan Zhou et al.ICML 2022 · 47 citations
- Learning Long-Term Reward Redistribution via Randomized Return DecompositionZhizhou Ren, Ruihan Guo, Yuan Zhou, Jian PengICLR 2022 · 45 citations
- Learning Energy Decompositions for Partial Inference in GFlowNetsHyosoon Jang, Minsu Kim, Sungsoo AhnICLR 2024 · 32 citations
Related papers
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang et al.ICML 2023 · 3 citations
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- DreamSmooth: Improving Model-based Reinforcement Learning via Reward SmoothingVint Lee, Pieter Abbeel, Youngwoon LeeICLR 2024 · 10 citations
- Milestone-Guided Policy Learning for Long-Horizon Language AgentsZixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan et al.ICML 2026 · 8 citations
