Augmenting Policy Learning with Routines Discovered from a Single Demonstration
Zelin Zhao, Chuang Gan, Jiajun Wu, Xiaoxiao Guo, Joshua B. Tenenbaum
摘要
Humans can abstract prior knowledge from very little data and use it to boost skill learning. In this paper, we propose routineaugmented policy learning (RAPL), which discovers routines composed of primitive actions from a single demonstration and uses discovered routines to augment policy learning. To discover routines from the demonstration, we first abstract routine candidates by identifying grammar over the demonstrated action trajectory. Then, the best routines measured by length and frequency are selected to form a routine library. We propose to learn policy simultaneously at primitive-level and routine-level with discovered routines, leveraging the temporal structure of routines. Our approach enables imitating expert behavior at multiple temporal scales for imitation learning and promotes reinforcement learning exploration. Extensive experiments on Atari games demonstrate that RAPL improves the state-of-the-art imitation learning method SQIL and reinforcement learning method A2C. Further, we show that discovered routines can generalize to unseen levels and difficulties on the CoinRun benchmark. * Work was done when Zelin was a visiting student at MIT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement LearningBoxuan Zhao, Jun Zhang, Deheng Ye, Jian Cao 等AAAI 2023 · 被引用 17 次
- Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD GenerationBo Yuan, Zelin Zhao, Petr Molodyk, Bin Hu 等ICML 2026 · 被引用 6 次
相关 Paper
- ASPiRe: Adaptive Skill Priors for Reinforcement LearningMengda Xu, Manuela Veloso, Shuran SongNeurIPS 2022 · 被引用 15 次
- Accelerating Robotic Reinforcement Learning via Parameterized Action PrimitivesMurtaza Dalal, Deepak Pathak, Ruslan SalakhutdinovNeurIPS 2021 · 被引用 121 次
- Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double ExplorationHeyang Zhao, Xingrui Yu, David Mark Bossens, Ivor W. Tsang 等ICLR 2025
- LISA: Learning Interpretable Skill Abstractions from LanguageDivyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song 等NeurIPS 2022 · 被引用 43 次
- Language-guided Skill Learning with Temporal Variational InferenceHaotian Fu, Pratyusha Sharma, Elias Stengel-Eskin, George Konidaris 等ICML 2024 · 被引用 11 次
