Generalized Behavior Learning from Diverse Demonstrations
Varshith Sreeramdass, Rohan R. Paleja, Letian Chen, Sanne van Waveren, Matthew C. Gombolay
摘要
Learning robot control policies through Reinforcement Learning can be challenging due to the complexity of designing rewards, which often result in unexpected behaviors. Imitation Learning overcomes this issue by using demonstrations to create policies that mimic expert behaviors. However, experts often demonstrate varied approaches to tasks. Capturing this variability is crucial for understanding and adapting to diverse scenarios. Prior methods capture variability by optimizing for behavior diversity alongside imitation. Yet, naive formulations of diversity can result in meaningless representation of latent factors, hindering generalization to novel scenarios. We propose Guided Strategy Discovery (GSD), a novel regularization method that specifically promotes expert-specified, taskrelevant diversity. In the recovery of unseen expert behaviors, GSD improves 11% over the next best baseline across three continuous control tasks on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Prompting Decision Transformer for Few-Shot Policy GeneralizationMengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu 等ICML 2022 · 被引用 194 次
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- Lipschitz-constrained Unsupervised Skill DiscoverySeohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee 等ICLR 2022 · 被引用 72 次
相关 Paper
- Open-Ended Diverse Solution Discovery with Regulated Behavior Patterns for Cross-Domain AdaptationKang Xu, Yan Ma, Bingsheng Wei, Wei LiAAAI 2023 · 被引用 3 次
- DGPO: Discovering Multiple Strategies with Diversity-Guided Policy OptimizationWentse Chen, Shiyu Huang, Yuan Chiang, Tim Pearce 等AAAI 2024 · 被引用 9 次
- Behavioral Mode Discovery for Fine-tuning Multimodal Generative PoliciesAlberta Longhini, David Emukpere, Jean-Michel Renders, Seungsu KimICML 2026 · 被引用 1 次
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement LearningJulien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai 等NeurIPS 2020 · 被引用 25 次
- Variational Imitation Learning with Diverse-quality DemonstrationsVoot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi SugiyamaICML 2020 · 被引用 38 次
