Offline Meta-Reinforcement Learning with Advantage Weighting
Eric Mitchell, Rafael Rafailov, Xue Bin Peng, Sergey Levine, Chelsea Finn
摘要
This paper introduces the offline metareinforcement learning (offline meta-RL) problem setting and proposes an algorithm that performs well in this setting. Offline meta-RL is analogous to the widely successful supervised learning strategy of pre-training a model on a large batch of fixed, pre-collected data (possibly from various tasks) and fine-tuning the model to a new task with relatively little data. That is, in offline meta-RL, we meta-train on fixed, pre-collected data from several tasks in order to adapt to a new task with a very small amount (less than 5 trajectories) of data from the new task. By nature of being offline, algorithms for offline meta-RL can utilize the largest possible pool of training data available and eliminate potentially unsafe or costly data collection during meta-training. This setting inherits the challenges of offline RL, but it differs significantly because offline RL does not generally consider a) transfer to new tasks or b) limited data from the test task, both of which we face in offline meta-RL. Targeting the offline meta-RL setting, we propose Meta-Actor Critic with Advantage Weighting (MACAW), an optimization-based meta-learning algorithm that uses simple, supervised regression objectives for both the inner and outer loop of meta-training. On offline variants of common meta-RL benchmarks, we empirically find that this approach enables fully offline meta-reinforcement learning and achieves notable gains over prior methods. Code available at https://sites.google.com/view/macaw-metarl .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- Prompting Decision Transformer for Few-Shot Policy GeneralizationMengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu 等ICML 2022 · 被引用 194 次
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak 等NeurIPS 2023 · 被引用 170 次
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 被引用 125 次
- Elastic Decision TransformerYueh-Hua Wu, Xiaolong Wang, Masashi HamayaNeurIPS 2023 · 被引用 96 次
- How to Leverage Unlabeled Data in Offline Reinforcement LearningTianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman 等ICML 2022 · 被引用 78 次
它引用的顶会 Paper5
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 被引用 162 次
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 被引用 132 次
- Information-theoretic Task Selection for Meta-Reinforcement LearningRicardo Luna Gutiérrez, Matteo LeonettiNeurIPS 2020 · 被引用 24 次
相关 Paper
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
- Model-Based Offline Meta-Reinforcement Learning with RegularizationSen Lin, Jialin Wan, Tengyu Xu, Yingbin Liang 等ICLR 2022 · 被引用 20 次
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 被引用 126 次
- Offline Multitask Representation Learning for Reinforcement LearningHaque Ishfaq, Thanh Nguyen-Tang, Songtao Feng, Raman Arora 等NeurIPS 2024 · 被引用 15 次
