Reinforcement Learning with Random Delays
Yann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal, Jonathan Binas
Abstract
Action and observation delays commonly occur in many Reinforcement Learning applications, such as remote control scenarios. We study the anatomy of randomly delayed environments, and show that partially resampling trajectory fragments in hindsight allows for off-policy multi-step value estimation. We apply this principle to derive Delay-Correcting Actor-Critic (DCAC), an algorithm based on Soft Actor-Critic with significantly better performance in environments with delays. This is shown theoretically and also demonstrated practically on a delay-augmented version of the MuJoCo continuous control benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Handling Delay in Real-Time Reinforcement LearningIvan Anokhin, Rishav Rishav, Matthew Riemer, Stephen Chung et al.ICLR 2025 · 225 citations
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 74 citations
- Learning Long-Term Reward Redistribution via Randomized Return DecompositionZhizhou Ren, Ruihan Guo, Yuan Zhou, Jian PengICLR 2022 · 45 citations
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 22 citations
- Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackJangwon Kim, Hangyeol Kim, Jiwook Kang, Jongchan Baek et al.NeurIPS 2023 · 16 citations
Builds on1
Related papers
- RVI-SAC: Average Reward Off-Policy Deep Reinforcement LearningYukinari Hisaki, Isao OnoICML 2024 · 6 citations
- Addressing Signal Delay in Deep Reinforcement LearningWilliam Wei Wang, Dongqi Han, Xufang Luo, Dongsheng LiICLR 2024 · 13 citations
- Continuous Soft Actor-Critic: An Off-Policy Learning Method Robust to Time DiscretizationHuimin Han, Shaolin JiNeurIPS 2025
- Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang et al.ICML 2024 · 9 citations
- Skill or Luck? Return Decomposition via Advantage FunctionsHsiao-Ru Pan, Bernhard SchölkopfICLR 2024 · 7 citations
