Delayed Reinforcement Learning by Imitation
Pierre Liotet, Davide Maran, Lorenzo Bisi, Marcello Restelli
摘要
When the agent’s observations or interactions are delayed, classic reinforcement learning tools usually fail. In this paper, we propose a simple yet new and efficient solution to this problem. We assume that, in the undelayed environment, an efficient policy is known or can be easily learned, but the task may suffer from delays in practice and we thus want to take them into account. We present a novel algorithm, Delayed Imitation with Dataset Aggregation (DIDA), which builds upon imitation learning methods to learn how to act in a delayed environment from undelayed demonstrations. We provide a theoretical analysis of the approach that will guide the practical design of DIDA. These results are also of general interest in the delayed reinforcement learning literature by providing bounds on the performance between delayed and undelayed tasks, under smoothness conditions. We show empirically that DIDA obtains high performances with a remarkable sample efficiency on a variety of tasks, including robotic locomotion, classic control, and trading.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Variational Delayed Policy OptimizationQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang 等NeurIPS 2024 · 被引用 10 次
- Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysQingyuan Wu, Simon Sinong Zhan, Yixuan Wang, Yuhui Wang 等ICML 2024 · 被引用 9 次
- No-Regret Reinforcement Learning in Smooth MDPsDavide Maran, Alberto Maria Metelli, Matteo Papini, Marcello RestelliICML 2024 · 被引用 6 次
- Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPsDavide Maran, Alberto Maria Metelli, Matteo Papini, Marcello RestelliNeurIPS 2024 · 被引用 6 次
- Rainbow Delay Compensation: A Multi-Agent Reinforcement Learning Framework for Mitigating Observation DelaysSongchen Fu, Siang Chen, Shaojing Zhao, Letian Bai 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper4
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 被引用 141 次
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement LearningAlberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi, Luca Sabbioni 等ICML 2020 · 被引用 43 次
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 被引用 4 次
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal 等ICLR 2021 · 被引用 3 次
相关 Paper
- Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouAAAI 2022 · 被引用 14 次
- Deterministic and Discriminative Imitation (D2-Imitation): Revisiting Adversarial Imitation for Sample EfficiencyMingfei Sun, Sam Devlin, Katja Hofmann, Shimon WhitesonAAAI 2022 · 被引用 7 次
- Efficient RL with Impaired Observability: Learning to Act with Delayed and Missing State ObservationsMinshuo Chen, Yu Bai, H. Vincent Poor, Mengdi WangNeurIPS 2023 · 被引用 19 次
- Addressing Signal Delay in Deep Reinforcement LearningWilliam Wei Wang, Dongqi Han, Xufang Luo, Dongsheng LiICLR 2024 · 被引用 13 次
- Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double ExplorationHeyang Zhao, Xingrui Yu, David Mark Bossens, Ivor W. Tsang 等ICLR 2025
