The Provable Benefit of Unsupervised Data Sharing for Offline Reinforcement Learning
Hao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie Zhang
Abstract
Self-supervised methods have become crucial for advancing deep learning by leveraging data itself to reduce the need for expensive annotations. However, the question of how to conduct self-supervised offline reinforcement learning (RL) in a principled way remains unclear. In this paper, we address this issue by investigating the theoretical benefits of utilizing reward-free data in linear Markov Decision Processes (MDPs) within a semi-supervised setting. Further, we propose a novel, Provable Data Sharing algorithm (PDS) to utilize such reward-free data for offline RL. PDS uses additional penalties on the reward function learned from labeled data to prevent overestimation, ensuring a conservative algorithm. Our results on various offline RL tasks demonstrate that PDS significantly improves the performance of offline RL algorithms with reward-free data. Overall, our work provides a promising approach to leveraging the benefits of unlabeled data in offline RL while maintaining theoretical guarantees. We believe our findings will contribute to developing more robust self-supervised RL methods. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsUladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero et al.NeurIPS 2025 · 109 citations
- Offline Imitation Learning with Model-based Reverse AugmentationJie-Jing Shao, Hao-Sen Shi, Lan-Zhe Guo, Yu-Feng LiKDD 2024 · 5 citations
- In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement LearningSongjun Tu, Jingbo Sun, Qichao Zhang, Yaocheng Zhang et al.AAAI 2025 · 4 citations
- Planning, Fast and Slow: Online Reinforcement Learning with Action-Free Offline Data via Multiscale PlannersChengjie Wu, Hao Hu, Yiqin Yang, Ning Zhang et al.ICML 2024 · 3 citations
- How to Solve Contextual Goal-Oriented Problems with Offline Datasets?Ying Fan, Jingling Li, Adith Swaminathan, Aditya Modi et al.NeurIPS 2024 · 1 citation
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- How to Leverage Unlabeled Data in Offline Reinforcement LearningTianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman et al.ICML 2022 · 78 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov GameWei Xiong, Han Zhong, Chengshuai Shi, Cong Shen et al.ICLR 2023 · 2 citations
- Semi-Supervised Offline Reinforcement Learning with Action-Free TrajectoriesQinqing Zheng, Mikael Henaff, Brandon Amos, Aditya GroverICML 2023 · 29 citations
- Offline Multitask Representation Learning for Reinforcement LearningHaque Ishfaq, Thanh Nguyen-Tang, Songtao Feng, Raman Arora et al.NeurIPS 2024 · 15 citations
