Imitation Learning from Vague Feedback
Xin-Qiang Cai, Yu-Jie Zhang, Chao-Kai Chiang, Masashi Sugiyama
摘要
Imitation learning from human feedback studies how to train well-performed imitation agents with an annotator’s relative comparison of two demonstrations (one demonstration is better/worse than the other), which is usually easier to collect than the perfect expert data required by traditional imitation learning. However, in many real-world applications, it is still expensive or even impossible to provide a clear pairwise comparison between two demonstrations with similar quality. This motivates us to study the problem of imitation learning with vague feedback, where the data annotator can only distinguish the paired demonstrations correctly when their quality differs significantly, i.e., one from the expert and another from the non-expert. By modeling the underlying demonstration pool as a mixture of expert and non-expert data, we show that the expert policy distribution can be recovered when the proportion α of expert data is known. We also propose a mixture proportion estimation method for the unknown α case. Then, we integrate the recovered expert policy distribution with generative adversarial imitation learning to form an end-to-end algorithm 1 . Experiments show that our methods outperform standard and preference-based imitation learning methods on various tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with ExplanationZelei Cheng, Xian Wu, Jiahao Yu, Sabrina Yang 等ICML 2024 · 被引用 11 次
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu 等ICML 2024 · 被引用 6 次
- Learning View-invariant World Models for Visual Robotic ManipulationJing-Cheng Pang, Nan Tang, Kaiyuan Li, Yuting Tang 等ICLR 2025
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Rethinking Importance Weighting for Deep Learning under Distribution ShiftTongtong Fang, Nan Lu, Gang Niu, Masashi SugiyamaNeurIPS 2020 · 被引用 179 次
- Mixture Proportion Estimation and PU Learning: A Modern ApproachSaurabh Garg, Yifan Wu, Alexander J. Smola, Sivaraman Balakrishnan 等NeurIPS 2021 · 被引用 79 次
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 被引用 73 次
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
相关 Paper
- Adversarial Imitation Learning with PreferencesAleksandar Taranovic, Andras Gabor Kupcsik, Niklas Freymuth, Gerhard NeumannICLR 2023 · 被引用 25 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
- Variational Imitation Learning with Diverse-quality DemonstrationsVoot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi SugiyamaICML 2020 · 被引用 38 次
- Unlabeled Imperfect Demonstrations in Adversarial Imitation LearningYunke Wang, Bo Du, Chang XuAAAI 2023 · 被引用 11 次
- f-GAIL: Learning f-Divergence for Generative Adversarial Imitation LearningXin Zhang, Yanhua Li, Ziming Zhang, Zhi-Li ZhangNeurIPS 2020 · 被引用 40 次
