Imitation Learning by Estimating Expertise of Demonstrators
Mark Beliaev, Andy Shih, Stefano Ermon, Dorsa Sadigh, Ramtin Pedarsani
摘要
Many existing imitation learning datasets are collected from multiple demonstrators, each with different expertise at different parts of the environment. Yet, standard imitation learning algorithms typically treat all demonstrators as homogeneous, regardless of their expertise, absorbing the weaknesses of any suboptimal demonstrators. In this work, we show that unsupervised learning over demonstrator expertise can lead to a consistent boost in the performance of imitation learning algorithms. We develop and optimize a joint model over a learned policy and expertise levels of the demonstrators. This enables our model to learn from the optimal behavior and filter out the suboptimal behavior of each demonstrator. Our model learns a single policy that can outperform even the best demonstrator, and can be used to estimate the expertise of any demonstrator at any state. We illustrate our findings on real-robotic continuous control tasks from Robomimic and discrete environments such as MiniGrid and chess, out-performing competing methods in out of settings, with an average of and up to improvement in terms of the final reward.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 被引用 135 次
- You Only Live Once: Single-Life Reinforcement LearningAnnie S. Chen, Archit Sharma, Sergey Levine, Chelsea FinnNeurIPS 2022 · 被引用 33 次
- Distance Weighted Supervised Learning for Offline Interaction DataJoey Hejna, Jensen Gao, Dorsa SadighICML 2023 · 被引用 21 次
- Blending Imitation and Reinforcement Learning for Robust Policy ImprovementXuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter 等ICLR 2024 · 被引用 19 次
- Selective Sampling and Imitation Learning via Online RegressionAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 被引用 15 次
它引用的顶会 Paper5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
- Aligning Superhuman AI with Human Behavior: Chess as a Model SystemReid McIlroy-Young, Siddhartha Sen, Jon M. Kleinberg, Ashton AndersonKDD 2020 · 被引用 77 次
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 被引用 73 次
- TRAIL: Near-Optimal Imitation Learning with Suboptimal DataMengjiao Yang, Sergey Levine, Ofir NachumICLR 2022 · 被引用 54 次
相关 Paper
- Variational Imitation Learning with Diverse-quality DemonstrationsVoot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi SugiyamaICML 2020 · 被引用 38 次
- Inverse Reinforcement Learning by Estimating Expertise of DemonstratorsMark Beliaev, Ramtin PedarsaniAAAI 2025 · 被引用 11 次
- Active Policy Improvement from Multiple Black-box OraclesXuefeng Liu, Takuma Yoneda, Chaoqi Wang, Matthew R. Walter 等ICML 2023 · 被引用 13 次
- Towards Uniformly Superhuman Autonomy via Subdominance MinimizationBrian D. Ziebart, Sanjiban Choudhury, Xinyan Yan, Paul VernazaICML 2022 · 被引用 2 次
- Learning Compound Tasks without Task-specific Knowledge via Imitation and Self-supervised LearningSang-Hyun Lee, Seung-Woo SeoICML 2020 · 被引用 12 次
