Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality
Songyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan Sui
Abstract
Most existing imitation learning approaches assume the demonstrations are drawn from experts who are optimal, but relaxing this assumption enables us to use a wider range of data. Standard imitation learning may learn a suboptimal policy from demonstrations with varying optimality. Prior works use confidence scores or rankings to capture beneficial information from demonstrations with varying optimality, but they suffer from many limitations, e.g., manually annotated confidence scores or high average optimality of demonstrations. In this paper, we propose a general framework to learn from demonstrations with varying optimality that jointly learns the confidence score and a well-performing policy. Our approach, Confidence-Aware Imitation Learning (CAIL) learns a well-performing policy from confidence-reweighted demonstrations, while using an outer loss to track the performance of our model and to learn the confidence. We provide theoretical guarantees on the convergence of CAIL and evaluate its performance in both simulated and real robot experiments. Our results show that CAIL significantly outperforms other imitation learning methods from demonstrations with varying optimality. We further show that even without access to any optimal demonstrations, CAIL can still learn a successful policy, and outperforms prior work. * Equal contribution. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 321be365-05d9-4edb-b5f9-549072ef3beaCited by top-tier papers25
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 105 citations
- Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement LearningRunze Liu, Fengshuo Bai, Yali Du, Yaodong YangNeurIPS 2022 · 72 citations
- Imitation Learning by Estimating Expertise of DemonstratorsMark Beliaev, Andy Shih, Stefano Ermon, Dorsa Sadigh et al.ICML 2022 · 60 citations
- Distance Weighted Supervised Learning for Offline Interaction DataJoey Hejna, Jensen Gao, Dorsa SadighICML 2023 · 21 citations
- Cross-Episodic Curriculum for Transformer AgentsLucy Xiaoyang Shi, Yunfan Jiang, Jake Grigsby, Linxi Fan et al.NeurIPS 2023 · 14 citations
Builds on2
Related papers
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu et al.ICML 2024 · 6 citations
- PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsQiang Liu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen et al.ICLR 2025
- Skill Disentanglement for Imitation Learning from Suboptimal DemonstrationsTianxiang Zhao, Wenchao Yu, Suhang Wang, Lu Wang et al.KDD 2023 · 5 citations
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 57 citations
- Inverse Reinforcement Learning by Estimating Expertise of DemonstratorsMark Beliaev, Ramtin PedarsaniAAAI 2025 · 11 citations
