Addressing Signal Delay in Deep Reinforcement Learning
William Wei Wang, Dongqi Han, Xufang Luo, Dongsheng Li
摘要
Despite the notable advancements in deep reinforcement learning (DRL) in recent years, a prevalent issue that is often overlooked is the impact of signal delay. Signal delay occurs when there is a lag between an agent's perception of the environment and its corresponding actions. In this paper, we first formalize delayed-observation Markov decision processes (DOMDP) by extending the standard MDP framework to incorporate signal delays. Next, we elucidate the challenges posed by the presence of signal delay in DRL, showing that trivial DRL algorithms and generic methods for partially observable tasks suffer greatly from delays. Lastly, we propose effective strategies to overcome these challenges. Our methods achieve remarkable performance in continuous robotic control tasks with large delays, yielding results comparable to those in non-delayed cases. Overall, our work contributes to a deeper understanding of DRL in the presence of signal delays and introduces novel approaches to address the associated challenges.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Variational Delayed Policy OptimizationQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang 等NeurIPS 2024 · 被引用 10 次
- Rainbow Delay Compensation: A Multi-Agent Reinforcement Learning Framework for Mitigating Observation DelaysSongchen Fu, Siang Chen, Shaojing Zhao, Letian Bai 等NeurIPS 2025 · 被引用 2 次
- Adaptive Reinforcement Learning for Unobservable Random DelaysJohn Wikman, Alexandre Proutiere, David BromanICML 2026
- Directly Forecasting Belief for Reinforcement Learning with DelaysQingyuan Wu, Yuhui Wang, Simon Sinong Zhan, Yixuan Wang 等ICML 2025
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 被引用 162 次
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 被引用 75 次
- Universal Trading for Order Execution with Oracle Policy DistillationYuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou 等AAAI 2021 · 被引用 52 次
- Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackJangwon Kim, Hangyeol Kim, Jiwook Kang, Jongchan Baek 等NeurIPS 2023 · 被引用 16 次
相关 Paper
- Thinking While Moving: Deep Reinforcement Learning with Concurrent ControlTed Xiao, Eric Jang, Dmitry Kalashnikov, Sergey Levine 等ICLR 2020 · 被引用 43 次
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 被引用 4 次
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal 等ICLR 2021 · 被引用 3 次
- Minimax Optimal Strategy for Delayed Observations in Online Reinforcement LearningHarin Lee, Kevin JamiesonICML 2026
- Efficient RL with Impaired Observability: Learning to Act with Delayed and Missing State ObservationsMinshuo Chen, Yu Bai, H. Vincent Poor, Mengdi WangNeurIPS 2023 · 被引用 19 次
