Variational Delayed Policy Optimization
Qingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang, Chung-Wei Lin, Chen Lv, Qi Zhu, Chao Huang
Abstract
In environments with delayed observation, state augmentation by including actions within the delay window is adopted to retrieve Markovian property to enable reinforcement learning (RL). However, state-of-the-art (SOTA) RL techniques with Temporal-Difference (TD) learning frameworks often suffer from learning inefficiency, due to the significant expansion of the augmented state space with the delay. To improve learning efficiency without sacrificing performance, this work introduces a novel framework called Variational Delayed Policy Optimization (VDPO), which reformulates delayed RL as a variational inference problem. This problem is further modelled as a two-step iterative optimization problem, where the first step is TD learning in the delay-free environment with a small state space, and the second step is behaviour cloning which can be addressed much more efficiently than TD learning. We not only provide a theoretical analysis of VDPO in terms of sample complexity and performance, but also empirically demonstrate that VDPO can achieve consistent performance with SOTA methods, with a significant enhancement of sample efficiency (approximately 50% less amount of samples) in the MuJoCo benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4dc1c34-ccbe-41a4-af60-477302671f26Cited by top-tier papers4
- Belief-Based Offline Reinforcement Learning for Delay-Robust Policy OptimizationSimon Sinong Zhan, Qingyuan Wu, Philip Wang, Frank Yang et al.ICLR 2026 · 1 citation
- TACTIC: Task-Aware Sparse Coordination Graphs for Multi-Task Multi-agent Reinforcement LearningKexing Peng, Pengyi Li, tinghuai ma, Jianye HaoICML 2026
- Adaptive Reinforcement Learning for Unobservable Random DelaysJohn Wikman, Alexandre Proutiere, David BromanICML 2026
- Directly Forecasting Belief for Reinforcement Learning with DelaysQingyuan Wu, Yuhui Wang, Simon Sinong Zhan, Yixuan Wang et al.ICML 2025
Builds on8
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu et al.ICML 2022 · 112 citations
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang et al.ICML 2023 · 81 citations
- Delayed Reinforcement Learning by ImitationPierre Liotet, Davide Maran, Lorenzo Bisi, Marcello RestelliICML 2022 · 22 citations
Related papers
- Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang et al.ICML 2024 · 9 citations
- Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit FeedbackTal Lancewicki, Aviv Rosenberg, Dmitry SotnikovICML 2023 · 6 citations
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 4 citations
- Policy Optimization with Stochastic Mirror DescentLong Yang, Yu Zhang, Gang Zheng, Qian Zheng et al.AAAI 2022 · 38 citations
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
