Learning Long-Term Reward Redistribution via Randomized Return Decomposition
Zhizhou Ren, Ruihan Guo, Yuan Zhou, Jian Peng
Abstract
Many practical applications of reinforcement learning require agents to learn from sparse and delayed rewards. It challenges the ability of agents to attribute their actions to future outcomes. In this paper, we consider the problem formulation of episodic reinforcement learning with trajectory feedback. It refers to an extreme delay of reward signals, in which the agent can only obtain one reward signal at the end of each trajectory. A popular paradigm for this problem setting is learning with a designed auxiliary dense reward function, namely proxy reward, instead of sparse environmental signals. Based on this framework, this paper proposes a novel reward redistribution algorithm, randomized return decomposition (RRD), to learn a proxy reward function for episodic reinforcement learning. We establish a surrogate problem by Monte-Carlo sampling that scales up least-squares-based reward redistribution to long-horizon problems. We analyze our surrogate loss function by connection with existing methods in the literature, which illustrates the algorithmic properties of our approach. In experiments, we extensively evaluate our proposed method on a variety of benchmark tasks with episodic rewards and demonstrate substantial improvement over baseline algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bd2879d-414c-4356-9959-e02d6226fa56Cited by top-tier papers19
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Better Training of GFlowNets with Local Credit and Incomplete TrajectoriesLing Pan, Nikolay Malkin, Dinghuai Zhang, Yoshua BengioICML 2023 · 100 citations
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 77 citations
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 74 citations
- Learning Energy Decompositions for Partial Inference in GFlowNetsHyosoon Jang, Minsu Kim, Sungsoo AhnICLR 2024 · 32 citations
Builds on17
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- RNA Secondary Structure Prediction By Learning Unrolled AlgorithmsXinshi Chen, Yu Li, Ramzan Umarov, Xin Gao et al.ICLR 2020 · 134 citations
Related papers
- Episodic Return Decomposition by Difference of Implicitly Assigned Sub-trajectory RewardHaoxin Lin, Hongqiu Wu, Jiaji Zhang, Yihao Sun et al.AAAI 2024 · 3 citations
- Reward Redistribution via Gaussian Process Likelihood EstimationMinheng Xiao, Xian YuAAAI 2026
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 46 citations
- Interpretable Reward Redistribution in Reinforcement Learning: A Causal ApproachYudi Zhang, Yali Du, Biwei Huang, Ziyan Wang et al.NeurIPS 2023 · 32 citations
- Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RLCharles Packer, Pieter Abbeel, Joseph E. GonzalezNeurIPS 2021 · 22 citations
