Offline Data Enhanced On-Policy Policy Gradient with Provable Guarantees
Yifei Zhou, Ayush Sekhari, Yuda Song, Wen Sun
Abstract
Hybrid RL is the setting where an RL agent has access to both offline data and online data by interacting with the real-world environment. In this work, we propose a new hybrid RL algorithm that combines an on-policy actor-critic method with offline data. On-policy methods such as policy gradient and natural policy gradient (NPG) have shown to be more robust to model misspecification, though sometimes it may not be as sample efficient as methods that rely on off-policy learning. On the other hand, offline methods that depend on off-policy training often require strong assumptions in theory and are less stable to train in practice. Our new approach integrates a procedure of off-policy training on the offline data into an on-policy NPG framework. We show that our approach, in theory, can obtain a best-of-both-worlds type of result -- it achieves the state-of-art theoretical guarantees of offline RL when offline RL-specific assumptions hold, while at the same time maintaining the theoretical guarantees of on-policy NPG regardless of the offline RL assumptions' validity. Experimentally, in challenging rich-observation environments, we show that our approach outperforms a state-of-the-art hybrid RL baseline which only relies on off-policy policy optimization, demonstrating the empirical benefit of combining on-policy and off-policy learning. Our code is publicly available at https://github.com/YifeiZhou02/HNPG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine et al.ICML 2024 · 163 citations
- REBEL: Reinforcement Learning via Regressing Relative RewardsZhaolin Gao, Jonathan D. Chang, Wenhao Zhan, Owen Oertell et al.NeurIPS 2024 · 82 citations
- Hybrid Inverse Reinforcement LearningJuntao Ren, Gokul Swamy, Steven Wu, Drew Bagnell et al.ICML 2024 · 33 citations
- Harnessing Density Ratios for Online Reinforcement LearningPhilip Amortila, Dylan J. Foster, Nan Jiang, Ayush Sekhari et al.ICLR 2024 · 14 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro et al.NeurIPS 2021 · 339 citations
Related papers
- Adaptive Policy Learning for Offline-to-Online Reinforcement LearningHan Zheng, Xufang Luo, Pengfei Wei, Xuan Song et al.AAAI 2023 · 47 citations
- Hybrid Reinforcement Learning from Offline Observation AloneYuda Song, Drew Bagnell, Aarti SinghICML 2024 · 6 citations
- Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement LearningGen Li, Wenhao Zhan, Jason D. Lee, Yuejie Chi et al.NeurIPS 2023 · 22 citations
- Hybrid RL: Using both offline and online data can make RL efficientYuda Song, Yifei Zhou, Ayush Sekhari, Drew Bagnell et al.ICLR 2023 · 7 citations
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
