Variational Imitation Learning with Diverse-quality Demonstrations
Voot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi Sugiyama
摘要
The goal of imitation learning (IL) is to learn a good policy from high-quality demonstrations. However, the quality of demonstrations in reality can be diverse, since it is easier and cheaper to collect demonstrations from a mix of experts and amateurs. IL in such situations can be challenging, especially when the level of demonstrators' expertise is unknown. We propose a new IL method called variational imitation learning with diverse-quality demonstrations (VILD), where we explicitly model the level of demonstrators' expertise with a probabilistic graphical model and estimate it along with a reward function. We show that a naive approach to estimation is not suitable to large state and action spaces, and fix its issues by using a variational approach which can be easily implemented using existing reinforcement learning methods. Experiments on continuous-control benchmarks demonstrate that VILD outperforms state-of-the-art methods. Our work enables scalable and data-efficient IL under more realistic settings than before.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 被引用 73 次
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingLantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger 等AAAI 2023 · 被引用 29 次
- Towards Robust Offline Reinforcement Learning under Diverse Data CorruptionRui Yang, Han Zhong, Jiawei Xu, Amy Zhang 等ICLR 2024 · 被引用 28 次
- Imitation Learning from Imperfection: Theoretical Justifications and AlgorithmsZiniu Li, Tian Xu, Zeyu Qin, Yang Yu 等NeurIPS 2023 · 被引用 26 次
相关 Paper
- Inverse Reinforcement Learning by Estimating Expertise of DemonstratorsMark Beliaev, Ramtin PedarsaniAAAI 2025 · 被引用 11 次
- Imitation Learning by Estimating Expertise of DemonstratorsMark Beliaev, Andy Shih, Stefano Ermon, Dorsa Sadigh 等ICML 2022 · 被引用 60 次
- Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline DataShilong Deng, Zetao Zheng, Hongcai He, Paul Weng 等AAAI 2025
- Imitation Learning from Vague FeedbackXin-Qiang Cai, Yu-Jie Zhang, Chao-Kai Chiang, Masashi SugiyamaNeurIPS 2023 · 被引用 5 次
- Visual Imitation Learning with Patch RewardsMinghuan Liu, Tairan He, Weinan Zhang, Shuicheng Yan 等ICLR 2023 · 被引用 1 次
