Episodic Return Decomposition by Difference of Implicitly Assigned Sub-trajectory Reward
Haoxin Lin, Hongqiu Wu, Jiaji Zhang, Yihao Sun, Junyin Ye, Yang Yu
摘要
Real-world decision-making problems are usually accompanied by delayed rewards, which affects the sample efficiency of Reinforcement Learning, especially in the extremely delayed case where the only feedback is the episodic reward obtained at the end of an episode. Episodic return decomposition is a promising way to deal with the episodic-reward setting. Several corresponding algorithms have shown remarkable effectiveness of the learned step-wise proxy rewards from return decomposition. However, these existing methods lack either attribution or representation capacity, leading to inefficient decomposition in the case of long-term episodes. In this paper, we propose a novel episodic return decomposition method called Diaster (Difference of implicitly assigned sub-trajectory reward). Diaster decomposes any episodic reward into credits of two divided sub-trajectories at any cut point, and the step-wise proxy rewards come from differences in expectation. We theoretically and empirically verify that the decomposed proxy reward function can guide the policy to be nearly optimal. Experimental results show that our method outperforms previous state-of-the-art methods in terms of both sample efficiency and performance. The code is available at https://github.com/HxLyn3/Diaster.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement LearningYun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao 等AAAI 2025 · 被引用 29 次
- Gradient-Guided Credit Assignment and Joint Optimization for Dependency-Aware Spatial CrowdsourcingYafei Li, Wei Chen, Jinxing Yan, Huiling Li 等AAAI 2025 · 被引用 3 次
- Reward Redistribution via Gaussian Process Likelihood EstimationMinheng Xiao, Xian YuAAAI 2026
它引用的顶会 Paper7
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 被引用 151 次
- Reinforcement Learning with Trajectory FeedbackYonathan Efroni, Nadav Merlis, Shie MannorAAAI 2021 · 被引用 48 次
- Off-Policy Reinforcement Learning with Delayed RewardsBeining Han, Zhizhou Ren, Zuofan Wu, Yuan Zhou 等ICML 2022 · 被引用 47 次
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 被引用 46 次
- Learning Long-Term Reward Redistribution via Randomized Return DecompositionZhizhou Ren, Ruihan Guo, Yuan Zhou, Jian PengICLR 2022 · 被引用 45 次
相关 Paper
- Interpretable Reward Redistribution in Reinforcement Learning: A Causal ApproachYudi Zhang, Yali Du, Biwei Huang, Ziyan Wang 等NeurIPS 2023 · 被引用 32 次
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 被引用 22 次
- RD: Reward Decomposition with Representation DecompositionZichuan Lin, Derek Yang, Li Zhao, Tao Qin 等NeurIPS 2020 · 被引用 12 次
- STAS: Spatial-Temporal Return Decomposition for Solving Sparse Rewards Problems in Multi-agent Reinforcement LearningSirui Chen, Zhaowei Zhang, Yaodong Yang, Yali DuAAAI 2024 · 被引用 11 次
- Skill or Luck? Return Decomposition via Advantage FunctionsHsiao-Ru Pan, Bernhard SchölkopfICLR 2024 · 被引用 7 次
