Composing Task-Agnostic Policies with Deep Reinforcement Learning
Ahmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson, Byron Boots, Michael C. Yip
摘要
The composition of elementary behaviors to solve challenging transfer learning problems is one of the key elements in building intelligent machines. To date, there has been plenty of work on learning task-specific policies or skills but almost no focus on composing necessary, task-agnostic skills to find a solution to new problems. In this paper, we propose a novel deep reinforcement learning-based skill transfer and composition method that takes the agent's primitive policies to solve unseen tasks. We evaluate our method in difficult cases where training policy through standard reinforcement learning (RL) or even hierarchical RL is either not feasible or exhibits high sample complexity. We show that our method not only transfers skills to new problem settings but also solves the challenging environments requiring both task planning and motion control with high data efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 被引用 247 次
- Learning to Coordinate Manipulation Skills via Skill Behavior DiversificationYoungwoon Lee, Jingyun Yang, Joseph J. LimICLR 2020 · 被引用 98 次
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- Lipschitz-constrained Unsupervised Skill DiscoverySeohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee 等ICLR 2022 · 被引用 72 次
- Controllability-Aware Unsupervised Skill DiscoverySeohong Park, Kimin Lee, Youngwoon Lee, Pieter AbbeelICML 2023 · 被引用 62 次
相关 Paper
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 被引用 23 次
- Skill Machines: Temporal Logic Skill Composition in Reinforcement LearningGeraud Nangue Tasse, Devon Jarvis, Steven James, Benjamin RosmanICLR 2024 · 被引用 12 次
- Toward Robust Long Range Policy TransferWei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, Min SunAAAI 2021 · 被引用 8 次
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 被引用 42 次
- Skill Discovery for Exploration and Planning using Deep Skill GraphsAkhil Bagaria, Jason K. Senthil, George KonidarisICML 2021 · 被引用 73 次
