Toward Robust Long Range Policy Transfer
Wei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, Min Sun
摘要
Humans can master a new task within a few trials by drawing upon skills acquired through prior experience. To mimic this capability, hierarchical models combining primitive policies learned from prior tasks have been proposed. However, these methods fall short comparing to the human's range of transferability. We propose a method, which leverages the hierarchical structure to train the combination function and adapt the set of diverse primitive polices alternatively, to efficiently produce a range of complex behaviors on challenging new tasks. We also design two regularization terms to improve the diversity and utilization rate of the primitives in the pre-training phase. We demonstrate that our method outperforms other recent policy transfer methods by combining and adapting these reusable primitives in tasks with continuous action space. The experiment results further show that our approach provides a broader transferring range. The ablation study also show the regularization terms are critical for long range policy transfer. Finally, we show that our method consistently outperforms other methods when the quality of the primitives varies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ASPiRe: Adaptive Skill Priors for Reinforcement LearningMengda Xu, Manuela Veloso, Shuran SongNeurIPS 2022 · 被引用 15 次
- Self-Composing Policies for Scalable Continual Reinforcement LearningMikel Malagón, Josu Ceberio, José Antonio LozanoICML 2024 · 被引用 13 次
- Flexible Attention-Based Multi-Policy Fusion for Efficient Deep Reinforcement LearningZih-Yun Chiu, Yi-Lin Tuan, William Yang Wang, Michael C. YipNeurIPS 2023 · 被引用 7 次
它引用的顶会 Paper3
- Sub-policy Adaptation for Hierarchical Reinforcement LearningAlexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter AbbeelICLR 2020 · 被引用 85 次
- Reinforcement Learning with Competitive Ensembles of Information-Constrained PrimitivesAnirudh Goyal, Shagun Sodhani, Jonathan Binas, Xue Bin Peng 等ICLR 2020 · 被引用 55 次
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson 等ICLR 2020 · 被引用 35 次
相关 Paper
- Hierarchically Decoupled Imitation For Morphological TransferDonald J. Hejna III, Lerrel Pinto, Pieter AbbeelICML 2020 · 被引用 47 次
- Meta-learning Parameterized SkillsHaotian Fu, Shangqun Yu, Saket Tiwari, Michael Littman 等ICML 2023 · 被引用 8 次
- Learning transferable motor skills with hierarchical latent mixture policiesDushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier 等ICLR 2022 · 被引用 34 次
- Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter MergingYajat Yadav, Zhiyuan Zhou, Andrew Wagenmaker, Karl Pertsch 等ICLR 2026 · 被引用 10 次
- CoMic: Complementary Task Learning & Mimicry for Reusable SkillsLeonard Hasenclever, Fabio Pardo, Raia Hadsell, Nicolas Heess 等ICML 2020 · 被引用 56 次
