What Matters for Batch Online Reinforcement Learning in Robotics?
Perry Dong, Suvir Mirchandani, Dorsa Sadigh, Chelsea Finn
摘要
The ability to learn from large batches of autonomously collected data for policy improvement---a paradigm we refer to as batch online reinforcement learning---holds the promise of enabling truly scalable robot learning by significantly reducing the need for human effort of data collection while getting benefits from self-improvement. Yet, despite the promise of this paradigm, it remains challenging to achieve due to algorithms not being able to learn effectively from the autonomous data. For example, prior works have applied imitation learning and filtered imitation learning methods to the batch online RL problem, but these algorithms often fail to efficiently improve from the autonomously collected data or converge quickly to a suboptimal point. This raises the question of what matters for effective batch online reinforcement learning in robotics. Motivated by this question, we perform a systematic empirical study of three axes---(i) algorithm class, (ii) policy extraction methods, and (iii) policy expressivity---and analyze how these axes affect performance and scaling with the amount of autonomously collected data. Through our analysis, we make several observations. First, we observe that the use of Q-functions to guide batch online RL significantly improves performance over imitation-based methods. Building on this, we show that an implicit method of policy extraction---via choosing the best action in the distribution of the policy---is necessary over traditional explicit policy extraction methods from offline RL. Next, we show that an expressive policy class is preferred over less expressive policy classes. Based on this analysis, we propose a general recipe for effective batch online RL. We then show a simple addition to the recipe, namely using temporally-correlated noise to obtain more diversity, results in further performance gains. Our recipe obtains significantly better performance and scaling compared to prior methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh 等ICLR 2026 · 被引用 13 次
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL FinetuningAndrew Wagenmaker, Perry Dong, Raymond Tsao, Chelsea Finn 等ICML 2026 · 被引用 10 次
- TQL: Scaling Q-Functions with Transformers by Preventing Attention CollapsePerry Dong, Kuo-Han Hung, Alexander Swerdlow, Dorsa Sadigh 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 被引用 173 次
相关 Paper
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 被引用 99 次
- Continuous Doubly Constrained Batch Reinforcement LearningRasool Fakoor, Jonas Mueller, Kavosh Asadi, Pratik Chaudhari 等NeurIPS 2021 · 被引用 37 次
- BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement LearningXinyue Chen, Zijian Zhou, Zheng Wang, Che Wang 等NeurIPS 2020 · 被引用 146 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
- Batch Reinforcement Learning with Hyperparameter GradientsByung-Jun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim 等ICML 2020 · 被引用 18 次
