DemoDICE: Offline Imitation Learning with Supplementary Imperfect Demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim
摘要
We consider offline imitation learning (IL), which aims to mimic the expert's behavior from its demonstration without further interaction with the environment. One of the main challenges in offline IL is to deal with the narrow support of the data distribution exhibited by the expert demonstrations that cover only a small fraction of the state and the action spaces. As a result, offline IL algorithms that rely only on expert demonstrations are very unstable since the situation easily deviates from those in the expert demonstrations. In this paper, we assume additional demonstration data of unknown degrees of optimality, which we call imperfect demonstrations. Under this setting, we propose DemoDICE, which effectively utilizes imperfect demonstrations by matching the stationary distribution of a policy with experts' distribution while penalizing its deviation from the overall demonstrations. Compared with the recent IL algorithms that adopt adversarial minimax training objectives, we substantially stabilize overall learning process by reducing minimax optimization to a direct convex optimization in a principled manner. Using extensive tasks, we show that DemoDICE achieves promising results in the offline IL from expert and imperfect demonstrations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper43
- Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement LearningLiyuan Mao, Haoran Xu, Xianyuan Zhan, Weinan Zhang 等NeurIPS 2024 · 被引用 49 次
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 被引用 48 次
- Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced DatasetsZhang-Wei Hong, Aviral Kumar, Sathwik Karnik, Abhishek Bhandwaldar 等NeurIPS 2023 · 被引用 34 次
- LobsDICE: Offline Learning from Observation via Stationary Distribution Correction EstimationGeon-Hyeong Kim, Jongmin Lee, Youngsoo Jang, Hongseok Yang 等NeurIPS 2022 · 被引用 33 次
- Survival Instinct in Offline Reinforcement LearningAnqi Li, Dipendra Misra, Andrey Kolobov, Ching-An ChengNeurIPS 2023 · 被引用 26 次
相关 Paper
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingLantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger 等AAAI 2023 · 被引用 29 次
- Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionSeokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang 等NeurIPS 2024 · 被引用 17 次
- Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal DatasetQuang Anh PHAM, Tien Mai, Akshat KumarICML 2026
- Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy MatchingYecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, Osbert BastaniICML 2022 · 被引用 49 次
- DualCOIL: Offline Imitation Learning from Contrasting DemonstrationsHuy Hoang, Tien Mai, Pradeep Varakantham, Tanvi VermaICML 2026
