Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
Qi Wang, Junming Yang, Yunbo Wang, Xin Jin, Wenjun Zeng, Xiaokang Yang
摘要
Training offline RL models using visual inputs poses two significant challenges, i.e., the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. This paper, in contrast, tries to build more flexible constraints for value estimation without impeding the exploration of potential advantages. The key idea is to leverage off-the-shelf RL simulators, which can be easily interacted with in an online manner, as the"test bed"for offline policies. To enable effective online-to-offline knowledge transfer, we introduce CoWorld, a model-based RL approach that mitigates cross-domain discrepancies in state and reward spaces. Experimental results demonstrate the effectiveness of CoWorld, outperforming existing RL approaches by large margins.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 被引用 15 次
- DreamSAC: Learning Hamiltonian World Models via Symmetry ExplorationJinzhou Tang, Fan Feng, Minghao Fu, Wenjun Lin 等CVPR 2026 · 被引用 1 次
- Video-Enhanced Offline Reinforcement Learning: A Model-Based ApproachMinting Pan, Yitao Zheng, Jiajian Li, Yunbo Wang 等ICML 2025
- Return-Critic: Bridging Goal Discrepancy for Efficient Visual Reinforcement LearningRuyi Lu, Xuesong Wang, Hengrui Zhang, Yuhu ChengICML 2026
- Open-World Reinforcement Learning over Long Short-Term ImaginationJiajian Li, Qi Wang, Yunbo Wang, Xin Jin 等ICLR 2025
它引用的顶会 Paper39
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
相关 Paper
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningQi Wang, Zhipeng Zhang, Baao Xie, Xin Jin 等ICCV 2025
- Compositional Conservatism: A Transductive Approach in Offline Reinforcement LearningYeda Song, Dongwook Lee, Gunhee KimICLR 2024 · 被引用 1 次
- Preference-based Policy Optimization from Sparse-reward Offline DatasetWenjie Qiu, Guofeng Cui, Shicheng Liu, Yuanlin Duan 等ICLR 2026
- Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangNeurIPS 2024 · 被引用 15 次
- Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningFuyuan Qian, Menglong Zhang, Song Wang, Quanying LiuICML 2026
