Imitation Learning by Estimating Expertise of Demonstrators
Mark Beliaev, Andy Shih, Stefano Ermon, Dorsa Sadigh, Ramtin Pedarsani
Abstract
Many existing imitation learning datasets are collected from multiple demonstrators, each with different expertise at different parts of the environment. Yet, standard imitation learning algorithms typically treat all demonstrators as homogeneous, regardless of their expertise, absorbing the weaknesses of any suboptimal demonstrators. In this work, we show that unsupervised learning over demonstrator expertise can lead to a consistent boost in the performance of imitation learning algorithms. We develop and optimize a joint model over a learned policy and expertise levels of the demonstrators. This enables our model to learn from the optimal behavior and filter out the suboptimal behavior of each demonstrator. Our model learns a single policy that can outperform even the best demonstrator, and can be used to estimate the expertise of any demonstrator at any state. We illustrate our findings on real-robotic continuous control tasks from Robomimic and discrete environments such as MiniGrid and chess, out-performing competing methods in out of settings, with an average of and up to improvement in terms of the final reward.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b91f05e-3ca7-495a-9b5e-f0b6a41dbf00Cited by top-tier papers14
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 135 citations
- You Only Live Once: Single-Life Reinforcement LearningAnnie S. Chen, Archit Sharma, Sergey Levine, Chelsea FinnNeurIPS 2022 · 33 citations
- Distance Weighted Supervised Learning for Offline Interaction DataJoey Hejna, Jensen Gao, Dorsa SadighICML 2023 · 21 citations
- Blending Imitation and Reinforcement Learning for Robust Policy ImprovementXuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter et al.ICLR 2024 · 19 citations
- Selective Sampling and Imitation Learning via Online RegressionAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 15 citations
Builds on5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 94 citations
- Aligning Superhuman AI with Human Behavior: Chess as a Model SystemReid McIlroy-Young, Siddhartha Sen, Jon M. Kleinberg, Ashton AndersonKDD 2020 · 77 citations
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 73 citations
- TRAIL: Near-Optimal Imitation Learning with Suboptimal DataMengjiao Yang, Sergey Levine, Ofir NachumICLR 2022 · 54 citations
Related papers
- Variational Imitation Learning with Diverse-quality DemonstrationsVoot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi SugiyamaICML 2020 · 38 citations
- Inverse Reinforcement Learning by Estimating Expertise of DemonstratorsMark Beliaev, Ramtin PedarsaniAAAI 2025 · 11 citations
- Active Policy Improvement from Multiple Black-box OraclesXuefeng Liu, Takuma Yoneda, Chaoqi Wang, Matthew R. Walter et al.ICML 2023 · 13 citations
- Towards Uniformly Superhuman Autonomy via Subdominance MinimizationBrian D. Ziebart, Sanjiban Choudhury, Xinyan Yan, Paul VernazaICML 2022 · 2 citations
- Learning Compound Tasks without Task-specific Knowledge via Imitation and Self-supervised LearningSang-Hyun Lee, Seung-Woo SeoICML 2020 · 12 citations
