Priors, Hierarchy, and Information Asymmetry for Skill Transfer in Reinforcement Learning
Sasha Salter, Kristian Hartikainen, Walter Goodwin, Ingmar Posner
摘要
The ability to discover behaviours from past experience and transfer them to new tasks is a hallmark of intelligent agents acting sample-efficiently in the real world. Equipping embodied reinforcement learners with the same ability may be crucial for their successful deployment in robotics. While hierarchical and KL-regularized reinforcement learning individually hold promise here, arguably a hybrid approach could combine their respective benefits. Key to these fields is the use of information asymmetry across architectural modules to bias which skills are learnt. While asymmetry choice has a large influence on transferability, existing methods base their choice primarily on intuition in a domain-independent, potentially sub-optimal, manner. In this paper, we theoretically and empirically show the crucial expressivity-transferability trade-off of skills across sequential tasks, controlled by information asymmetry. Given this insight, we introduce Attentive Priors for Expressive and Transferable Skills (APES), a hierarchical KL-regularized method, heavily benefiting from both priors and hierarchy. Unlike existing approaches, APES automates the choice of asymmetry by learning it in a data-driven, domain-dependent, way based on our expressivity-transferability theorems. Experiments over complex transfer domains of varying levels of extrapolation and sparsity, such as robot block stacking, demonstrate the criticality of the correct asymmetric choice, with APES drastically outperforming previous methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu 等ICLR 2021 · 被引用 161 次
- Catch & Carry: reusable neural controllers for vision-guided whole-body tasksJosh Merel, Saran Tunyasuvunakool, Arun Ahuja, Yuval Tassa 等SIGGRAPH 2020 · 被引用 103 次
- Hierarchical Skills for Efficient ExplorationJonas Gehring, Gabriel Synnaeve, Andreas Krause, Nicolas UsunierNeurIPS 2021 · 被引用 52 次
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon 等AAAI 2020 · 被引用 51 次
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe 等ICML 2021 · 被引用 48 次
相关 Paper
- ASPiRe: Adaptive Skill Priors for Reinforcement LearningMengda Xu, Manuela Veloso, Shuran SongNeurIPS 2022 · 被引用 15 次
- Hierarchically Decoupled Imitation For Morphological TransferDonald J. Hejna III, Lerrel Pinto, Pieter AbbeelICML 2020 · 被引用 47 次
- Toward Robust Long Range Policy TransferWei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, Min SunAAAI 2021 · 被引用 8 次
- Sub-policy Adaptation for Hierarchical Reinforcement LearningAlexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter AbbeelICLR 2020 · 被引用 85 次
- Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement LearningJinxin Liu, Hao Shen, Donglin Wang, Yachen Kang 等NeurIPS 2021 · 被引用 20 次
