Acting in Delayed Environments with Non-Stationary Markov Policies
Esther Derman, Gal Dalal, Shie Mannor
Abstract
The standard Markov Decision Process (MDP) formulation hinges on the assumption that an action is executed immediately after it was chosen. However, assuming it is often unrealistic and can lead to catastrophic failures in applications such as robotic manipulation, cloud computing, and finance. We introduce a framework for learning and planning in MDPs where the decision-maker commits actions that are executed with a delay of steps. The brute-force state augmentation baseline where the state is concatenated to the last committed actions suffers from an exponential complexity in , as we show for policy iteration. We then prove that with execution delay, deterministic Markov policies in the original state-space are sufficient for attaining maximal reward, but need to be non-stationary. As for stationary Markov policies, we show they are sub-optimal in general. Consequently, we devise a non-stationary Q-learning style model-based algorithm that solves delayed execution tasks without resorting to state-augmentation. Experiments on tabular, physical, and Atari domains reveal that it converges quickly to high performance even for substantial delays, while standard approaches that either ignore the delay or rely on state-augmentation struggle or fail due to divergence. The code is available at github.com/galdl/rl_delay_basic and github.com/galdl/rl_delay_atari.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d654b50a-19a7-44f5-8a14-a2d3d285877bCited by top-tier papers10
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 77 citations
- Near-Optimal Regret for Adversarial MDP with Delayed Bandit FeedbackTiancheng Jin, Tal Lancewicki, Haipeng Luo, Yishay Mansour et al.NeurIPS 2022 · 29 citations
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 22 citations
- Belief Projection-Based Reinforcement Learning for Environments with Delayed FeedbackJangwon Kim, Hangyeol Kim, Jiwook Kang, Jongchan Baek et al.NeurIPS 2023 · 16 citations
- Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang et al.ICML 2024 · 9 citations
Builds on2
Related papers
- Tree Search-Based Policy Optimization under Stochastic Execution DelayDavid Valensi, Esther Derman, Shie Mannor, Gal DalalICLR 2024 · 4 citations
- Minimax Optimal Strategy for Delayed Observations in Online Reinforcement LearningHarin Lee, Kevin JamiesonICML 2026
- Addressing Signal Delay in Deep Reinforcement LearningWilliam Wei Wang, Dongqi Han, Xufang Luo, Dongsheng LiICLR 2024 · 13 citations
- Variational Delayed Policy OptimizationQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang et al.NeurIPS 2024 · 10 citations
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 96 citations
