Offline Meta Reinforcement Learning - Identifiability Challenges and Effective Data Collection Strategies
Ron Dorfman, Idan Shenfeld, Aviv Tamar
摘要
Consider the following instance of the Offline Meta Reinforcement Learning (OMRL) problem: given the complete training logs of N conventional RL agents, trained on N different tasks, design a meta-agent that can quickly maximize reward in a new, unseen task from the same task distribution. In particular, while each conventional RL agent explored and exploited its own different task, the meta-agent must identify regularities in the data that lead to effective exploration/exploitation in the unseen task. Here, we take a Bayesian RL (BRL) view, and seek to learn a Bayes-optimal policy from the offline data. Building on the recent VariBAD BRL approach, we develop an off-policy BRL method that learns to plan an exploration strategy based on an adaptive neural belief estimate. However, learning to infer such a belief from offline data brings a new identifiability issue we term MDP ambiguity. We characterize the problem, and suggest resolutions via data collection and modification procedures. Finally, we evaluate our framework on a diverse set of domains, including difficult sparse reward tasks, and demonstrate learning of effective exploration behavior that is qualitatively different from the exploration used by any RL agent in the data. Our code is available online at https://github.com/Rondorf/BOReL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak 等NeurIPS 2023 · 被引用 170 次
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 被引用 80 次
- How to Leverage Unlabeled Data in Offline Reinforcement LearningTianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman 等ICML 2022 · 被引用 78 次
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised PretrainingLicong Lin, Yu Bai, Song MeiICLR 2024 · 被引用 74 次
- Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence ModelingSili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang 等NeurIPS 2024 · 被引用 65 次
它引用的顶会 Paper4
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Offline Meta-Reinforcement Learning with Advantage WeightingEric Mitchell, Rafael Rafailov, Xue Bin Peng, Sergey Levine 等ICML 2021 · 被引用 122 次
- Multi-task Batch Reinforcement Learning with Metric LearningJiachen Li, Quan Vuong, Shuang Liu, Minghua Liu 等NeurIPS 2020 · 被引用 64 次
相关 Paper
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
- Offline Meta Reinforcement Learning with In-Distribution Online AdaptationJianhao Wang, Jin Zhang, Haozhe Jiang, Junyu Zhang 等ICML 2023 · 被引用 16 次
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement LearningMenglong Zhang, Fuyuan Qian, Quanying LiuICLR 2025
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen 等ICML 2021 · 被引用 33 次
- Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian LensJihwan Jeong, Xiaoyu Wang, Jingmin Wang, Scott Sanner 等ICML 2025
