Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution Correction
Seokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim
摘要
We consider offline imitation learning (IL), which aims to train an agent to imitate from the dataset of expert demonstrations without online interaction with the environment. Behavioral Cloning (BC) has been a simple yet effective approach to offline IL, but it is also well-known to be vulnerable to the covariate shift resulting from the mismatch between the state distributions induced by the learned policy and the expert policy. Moreover, as often occurs in practice, when expert datasets are collected from an arbitrary state distribution instead of a stationary one, these shifts become more pronounced, potentially leading to substantial failures in existing IL methods. Specifically, we focus on covariate shift resulting from arbitrary state data distributions, such as biased data collection or incomplete trajectories, rather than shifts induced by changes in dynamics or noisy expert actions. In this paper, to mitigate the effect of the covariate shifts in BC, we propose DrilDICE, which utilizes a distributionally robust BC objective by employing a stationary distribution correction ratio estimation (DICE) to derive a feasible solution. We evaluate the effectiveness of our method through an extensive set of experiments covering diverse covariate shift scenarios. The results demonstrate the efficacy of the proposed approach in improving the robustness against the shifts, outperforming existing offline IL methods in such scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On Robustness of Vision-Language-Action Model against Multi-Modal PerturbationsJianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma 等ICLR 2026 · 被引用 7 次
- Flow Actor-Critic for Offline Reinforcement LearningJongseong Chae, Jongeui Park, Yongjae Shin, Gyeongmin Kim 等ICLR 2026 · 被引用 7 次
- VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise InstructionsHung-Ting Su, Ting-Jun Wang, Jia-Fong Yeh, Min Sun 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper12
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau 等ICML 2021 · 被引用 137 次
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon 等ICLR 2022 · 被引用 111 次
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 被引用 105 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
相关 Paper
- A Simple Solution for Offline Imitation from Observations and Examples with Possibly Incomplete TrajectoriesKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2023 · 被引用 7 次
- Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal DatasetQuang Anh PHAM, Tien Mai, Akshat KumarICML 2026
- ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient UpdateLiyuan Mao, Haoran Xu, Weinan Zhang, Xianyuan ZhanICLR 2024 · 被引用 23 次
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingLantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger 等AAAI 2023 · 被引用 29 次
- Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy OptimizationWoosung Kim, Donghyeon Ki, Byung-Jun LeeAAAI 2024 · 被引用 4 次
