Addressing Signal Delay in Deep Reinforcement Learning
William Wei Wang, Dongqi Han, Xufang Luo, Dongsheng Li
Abstract
Despite the notable advancements in deep reinforcement learning (DRL) in recent years, a prevalent issue that is often overlooked is the impact of signal delay. Signal delay occurs when there is a lag between an agent's perception of the environment and its corresponding actions. In this paper, we first formalize delayed-observation Markov decision processes (DOMDP) by extending the standard MDP framework to incorporate signal delays. Next, we elucidate the challenges posed by the presence of signal delay in DRL, showing that trivial DRL algorithms and generic methods for partially observable tasks suffer greatly from delays. Lastly, we propose effective strategies to overcome these challenges. Our methods achieve remarkable performance in continuous robotic control tasks with large delays, yielding results comparable to those in non-delayed cases. Overall, our work contributes to a deeper understanding of DRL in the presence of signal delays and introduces novel approaches to address the associated challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d014d77-c2c9-4cec-ad91-0e3458b1ae0dCited by top-tier papers4
- Variational Delayed Policy OptimizationQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang et al.NeurIPS 2024 · 10 citations
- Rainbow Delay Compensation: A Multi-Agent Reinforcement Learning Framework for Mitigating Observation DelaysSongchen Fu, Siang Chen, Shaojing Zhao, Letian Bai et al.NeurIPS 2025 · 2 citations
- Adaptive Reinforcement Learning for Unobservable Random DelaysJohn Wikman, Alexandre Proutiere, David BromanICML 2026
- Directly Forecasting Belief for Reinforcement Learning with DelaysQingyuan Wu, Yuhui Wang, Simon Sinong Zhan, Yixuan Wang et al.ICML 2025
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
- Universal Trading for Order Execution with Oracle Policy DistillationYuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou et al.AAAI 2021 · 52 citations
- Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackJangwon Kim, Hangyeol Kim, Jiwook Kang, Jongchan Baek et al.NeurIPS 2023 · 16 citations
Related papers
- Thinking While Moving: Deep Reinforcement Learning with Concurrent ControlTed Xiao, Eric Jang, Dmitry Kalashnikov, Sergey Levine et al.ICLR 2020 · 43 citations
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 4 citations
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal et al.ICLR 2021 · 3 citations
- Minimax Optimal Strategy for Delayed Observations in Online Reinforcement LearningHarin Lee, Kevin JamiesonICML 2026
- Efficient RL with Impaired Observability: Learning to Act with Delayed and Missing State ObservationsMinshuo Chen, Yu Bai, H. Vincent Poor, Mengdi WangNeurIPS 2023 · 19 citations
