Off-Policy Reinforcement Learning with Delayed Rewards
Beining Han, Zhizhou Ren, Zuofan Wu, Yuan Zhou, Jian Peng
摘要
We study deep reinforcement learning (RL) algorithms with delayed rewards. In many real-world tasks, instant rewards are often not readily accessible or even defined immediately after the agent performs actions. In this work, we first formally define the environment with delayed rewards and discuss the challenges raised due to the non-Markovian nature of such environments. Then, we introduce a general off-policy RL framework with a new Q-function formulation that can handle the delayed rewards with theoretical convergence guarantees. For practical tasks with high dimensional state spaces, we further introduce the HC-decomposition rule of the Q-function in our framework which naturally leads to an approximation scheme that helps boost the training efficiency and stability. We finally conduct extensive experiments to demonstrate the superior performance of our algorithms over the existing work and their variants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 被引用 74 次
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 被引用 60 次
- Learning Long-Term Reward Redistribution via Randomized Return DecompositionZhizhou Ren, Ruihan Guo, Yuan Zhou, Jian PengICLR 2022 · 被引用 45 次
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningRoger Creus Castanyer, Johan S. Obando-Ceron, Lu Li, Pierre-Luc Bacon 等NeurIPS 2025 · 被引用 26 次
- Off-Policy Evaluation for Human FeedbackQitong Gao, Ge Gao, Juncheng Dong, Vahid Tarokh 等NeurIPS 2023 · 被引用 13 次
它引用的顶会 Paper1
相关 Paper
- Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit FeedbackTal Lancewicki, Aviv Rosenberg, Dmitry SotnikovICML 2023 · 被引用 6 次
- Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionHongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen 等AAAI 2021 · 被引用 5 次
- Deep PQR: Solving Inverse Reinforcement Learning using Anchor ActionsSinong Geng, Houssam Nassif, Carlos A. Manzanares, A. Max Reppen 等ICML 2020 · 被引用 14 次
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 被引用 4 次
- Posterior Sampling with Delayed Feedback for Reinforcement Learning with Linear Function ApproximationNikki Lijing Kuang, Ming Yin, Mengdi Wang, Yu-Xiang Wang 等NeurIPS 2023 · 被引用 8 次
