Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
Simon Sinong Zhan, Qingyuan Wu, Philip Wang, Frank Yang, Xiangyu Shi, Chao Huang, Qi Zhu
摘要
Offline–to–online deployment of reinforcement learning (RL) agents often stumbles over two fundamental gaps: (1) the sim-to-real gap, where real-world systems exhibit latency and other physical imperfections not captured in simulation; and (2) the interaction gap, where policies trained purely offline face out-of-distribution (OOD) issues during online execution, as collecting new interaction data is costly or risky. As a result, agents must generalize from static, delay-free datasets to dynamic, delay-prone environments. In this work, we propose (elay-ransformer belief policy onstrained ffline ), a novel framework for learning delay-resilient policies solely from static, delay-free offline data. DT-CORL introduces a transformer-based belief model to infer latent states from delayed observations and jointly trains this belief with a constrained policy objective, ensuring that value estimation and belief representation remain aligned throughout learning. Crucially, our method does not require access to delayed transitions during training and outperforms naive history-augmented baselines, SOTA delayed RL methods, and existing belief-based approaches. Empirically, we demonstrate that DT-CORL achieves strong delay-robust generalization across both locomotion and goal-conditioned tasks in the D4RL benchmark under varying delay regimes. Our results highlight that joint belief-policy optimization is essential for bridging the sim-to-real latency gap and achieving stable performance in delayed environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
相关 Paper
- Directly Forecasting Belief for Reinforcement Learning with DelaysQingyuan Wu, Yuhui Wang, Simon Sinong Zhan, Yixuan Wang 等ICML 2025
- Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RLJunyu Guo, Zhi Zheng, Donghao Ying, Ming Jin 等NeurIPS 2025 · 被引用 2 次
- Decision Transformer under Random Frame DroppingKaizhe Hu, Ray Chen Zheng, Yang Gao, Huazhe XuICLR 2023 · 被引用 3 次
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li 等NeurIPS 2022 · 被引用 81 次
- Addressing Optimism Bias in Sequence Modeling for Reinforcement LearningAdam R. Villaflor, Zhe Huang, Swapnil Pande, John M. Dolan 等ICML 2022 · 被引用 30 次
