Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning
Gen Li, Wenhao Zhan, Jason D. Lee, Yuejie Chi, Yuxin Chen
摘要
This paper studies tabular reinforcement learning (RL) in the hybrid setting, which assumes access to both an offline dataset and online interactions with the unknown environment. A central question boils down to how to efficiently utilize online data collection to strengthen and complement the offline dataset and enable effective policy fine-tuning. Leveraging recent advances in reward-agnostic exploration and model-based offline RL, we design a three-stage hybrid RL algorithm that beats the best of both worlds -- pure offline RL and pure online RL -- in terms of sample complexities. The proposed algorithm does not require any reward information during data collection. Our theory is developed based on a new notion called single-policy partial concentrability, which captures the trade-off between distribution mismatch and miscoverage and guides the interplay between offline and online data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Harnessing Density Ratios for Online Reinforcement LearningPhilip Amortila, Dylan J. Foster, Nan Jiang, Ayush Sekhari 等ICLR 2024 · 被引用 14 次
- Scalable Online Exploration via CoverabilityPhilip Amortila, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 被引用 10 次
- Hybrid Reinforcement Learning Breaks Sample Size Barriers In Linear MDPsKevin Tan, Wei Fan, Yuting WeiNeurIPS 2024 · 被引用 6 次
- Hybrid Reinforcement Learning from Offline Observation AloneYuda Song, Drew Bagnell, Aarti SinghICML 2024 · 被引用 6 次
- On The Statistical Complexity of Offline Decision-MakingThanh Nguyen-Tang, Raman AroraICML 2024 · 被引用 2 次
它引用的顶会 Paper29
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro 等NeurIPS 2021 · 被引用 339 次
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
相关 Paper
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong 等NeurIPS 2021 · 被引用 207 次
- Hybrid RL: Using both offline and online data can make RL efficientYuda Song, Yifei Zhou, Ayush Sekhari, Drew Bagnell 等ICLR 2023 · 被引用 7 次
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
- Offline Data Enhanced On-Policy Policy Gradient with Provable GuaranteesYifei Zhou, Ayush Sekhari, Yuda Song, Wen SunICLR 2024 · 被引用 11 次
- Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline DataRuiqi Zhang, Andrea ZanetteNeurIPS 2023 · 被引用 12 次
