BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning
Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, Keith W. Ross
摘要
There has recently been a surge in research in batch Deep Reinforcement Learning (DRL), which aims for learning a high-performing policy from a given dataset without additional interactions with the environment. We propose a new algorithm, Best-Action Imitation Learning (BAIL), which strives for both simplicity and performance. BAIL learns a V function, uses the V function to select actions it believes to be high-performing, and then uses those actions to train a policy network using imitation learning. For the MuJoCo benchmark, we provide a comprehensive experimental study of BAIL, comparing its performance to four other batch Q-learning and imitation-learning schemes for a large variety of batch datasets. Our experiments show that BAIL's performance is much higher than the other schemes, and is also computationally much faster than the batch Q-learning schemes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
- Offline RL Without Off-Policy EvaluationDavid Brandfonbrener, Will Whitney, Rajesh Ranganath, Joan BrunaNeurIPS 2021 · 被引用 217 次
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 被引用 173 次
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 被引用 121 次
相关 Paper
- Continuous Doubly Constrained Batch Reinforcement LearningRasool Fakoor, Jonas Mueller, Kavosh Asadi, Pratik Chaudhari 等NeurIPS 2021 · 被引用 37 次
- What Matters for Batch Online Reinforcement Learning in Robotics?Perry Dong, Suvir Mirchandani, Dorsa Sadigh, Chelsea FinnICLR 2026 · 被引用 18 次
- Batch Reinforcement Learning with Hyperparameter GradientsByung-Jun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim 等ICML 2020 · 被引用 18 次
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu 等ICML 2024 · 被引用 6 次
- Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double ExplorationHeyang Zhao, Xingrui Yu, David Mark Bossens, Ivor W. Tsang 等ICLR 2025
