Adaptive Reinforcement Learning for Unobservable Random Delays
John Wikman, Alexandre Proutiere, David Broman
Abstract
In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instantaneously, selects an action without delay, and executes it immediately. In real-world dynamic environments, such as cyber-physical systems, this assumption often breaks down due to delays in the interaction between the agent and the system. These delays can vary stochastically over time and are typically unobservable when deciding on an action. Existing methods deal with this uncertainty conservatively by assuming a known fixed upper bound on the delay, even if the delay is often much lower. In this work, we introduce the interaction layer , a general framework that enables agents to adaptively handle unobservable and time-varying delays. Specifically, the agent generates a matrix of possible future actions, anticipating a horizon of potential delays, to handle both unpredictable delays and lost action packets sent over networks. Building on this framework, we develop a model-based algorithm, Actor-Critic with Delay Adaptation (ACDA) , which dynamically adjusts to delay patterns. Our method significantly outperforms state-of-the-art approaches across a wide range of locomotion benchmark environments, including real-world measured delays.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8bf7444-96f2-4ba4-8947-e05c899ce55aBuilds on9
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 114 citations
- Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackJangwon Kim, Hangyeol Kim, Jiwook Kang, Jongchan Baek et al.NeurIPS 2023 · 16 citations
- Addressing Signal Delay in Deep Reinforcement LearningWilliam Wei Wang, Dongqi Han, Xufang Luo, Dongsheng LiICLR 2024 · 13 citations
- Variational Delayed Policy OptimizationQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang et al.NeurIPS 2024 · 10 citations
Related papers
- Learn to change the world: Multi-level reinforcement learning with model-changing actionsZiqing Lu, Babak Hassibi, Lifeng Lai, Weiyu XuICML 2026
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal et al.ICLR 2021 · 3 citations
- Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable SimulationIgnat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden et al.ICML 2024 · 27 citations
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 22 citations
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler et al.NeurIPS 2024 · 10 citations
