Reinforcement Learning with Random Delays
Yann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal, Jonathan Binas
摘要
Action and observation delays commonly occur in many Reinforcement Learning applications, such as remote control scenarios. We study the anatomy of randomly delayed environments, and show that partially resampling trajectory fragments in hindsight allows for off-policy multi-step value estimation. We apply this principle to derive Delay-Correcting Actor-Critic (DCAC), an algorithm based on Soft Actor-Critic with significantly better performance in environments with delays. This is shown theoretically and also demonstrated practically on a delay-augmented version of the MuJoCo continuous control benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Handling Delay in Real-Time Reinforcement LearningIvan Anokhin, Rishav Rishav, Matthew Riemer, Stephen Chung 等ICLR 2025 · 被引用 225 次
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 被引用 74 次
- Learning Long-Term Reward Redistribution via Randomized Return DecompositionZhizhou Ren, Ruihan Guo, Yuan Zhou, Jian PengICLR 2022 · 被引用 45 次
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 被引用 22 次
- Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackJangwon Kim, Hangyeol Kim, Jiwook Kang, Jongchan Baek 等NeurIPS 2023 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- RVI-SAC: Average Reward Off-Policy Deep Reinforcement LearningYukinari Hisaki, Isao OnoICML 2024 · 被引用 6 次
- Addressing Signal Delay in Deep Reinforcement LearningWilliam Wei Wang, Dongqi Han, Xufang Luo, Dongsheng LiICLR 2024 · 被引用 13 次
- Continuous Soft Actor-Critic: An Off-Policy Learning Method Robust to Time DiscretizationHuimin Han, Shaolin JiNeurIPS 2025
- Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang 等ICML 2024 · 被引用 9 次
- Skill or Luck? Return Decomposition via Advantage FunctionsHsiao-Ru Pan, Bernhard SchölkopfICLR 2024 · 被引用 7 次
