Meta Reinforcement Learning with Autonomous Inference of Subtask Dependencies
Sungryull Sohn, Hyunjae Woo, Jongwook Choi, Honglak Lee
摘要
We propose and address a novel few-shot RL problem, where a task is characterized by a subtask graph which describes a set of subtasks and their dependencies that are unknown to the agent. The agent needs to quickly adapt to the task over few episodes during adaptation phase to maximize the return in the test phase. Instead of directly learning a meta-policy, we develop a Meta-learner with Subtask Graph Inference(MSGI), which infers the latent parameter of the task by interacting with the environment and maximizes the return given the latent parameter. To facilitate learning, we adopt an intrinsic reward inspired by upper confidence bound (UCB) that encourages efficient exploration. Our experiment results on two grid-world domains and StarCraft II environments show that the proposed method is able to accurately infer the latent task parameter, and to adapt more efficiently than existing meta RL and hierarchical RL methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Environment Generation for Zero-Shot Compositional Reinforcement LearningIzzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi 等NeurIPS 2021 · 被引用 50 次
- Recomposing the Reinforcement Learning Building Blocks with HypernetworksElad Sarafian, Shai Keynan, Sarit KrausICML 2021 · 被引用 42 次
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 被引用 36 次
- Discovering Hierarchical Achievements in Reinforcement Learning via Contrastive LearningSeungyong Moon, Junyoung Yeom, Bumsoo Park, Hyun Oh SongNeurIPS 2023 · 被引用 12 次
- Learning Parameterized Task Structure for Generalization to Unseen EntitiesAnthony Z. Liu, Sungryull Sohn, Mahdi Qazwini, Honglak LeeAAAI 2022 · 被引用 6 次
相关 Paper
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen 等ICML 2021 · 被引用 33 次
- Efficient Meta Reinforcement Learning for Preference-based Fast AdaptationZhizhou Ren, Anji Liu, Yitao Liang, Jian Peng 等NeurIPS 2022 · 被引用 11 次
- Graph Meta Learning via Local SubgraphsKexin Huang, Marinka ZitnikNeurIPS 2020 · 被引用 205 次
- Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and PlanningYizhe Huang, Anji Liu, Fanqi Kong, Yaodong Yang 等ICML 2024 · 被引用 5 次
- Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationChristopher R. Dance, Julien Perez, Théo CachetICML 2021 · 被引用 17 次
