Offline Imitation Learning with Variational Counterfactual Reasoning
Zexu Sun, Bowei He, Jinxin Liu, Xu Chen, Chen Ma, Shuai Zhang
摘要
In offline imitation learning (IL), an agent aims to learn an optimal expert behavior policy without additional online environment interactions. However, in many real-world scenarios, such as robotics manipulation, the offline dataset is collected from suboptimal behaviors without rewards. Due to the scarce expert data, the agents usually suffer from simply memorizing poor trajectories and are vulnerable to variations in the environments, lacking the capability of generalizing to new environments. To automatically generate high-quality expert data and improve the generalization ability of the agent, we propose a framework named Offline Imitation Learning with Counterfactual data Augmentation (OILCA) by doing counterfactual inference. In particular, we leverage identifiable variational autoencoder to generate counterfactual samples for expert data augmentation. We theoretically analyze the influence of the generated expert data and the improvement of generalization. Moreover, we conduct extensive experiments to demonstrate that our approach significantly outperforms various baselines on both DeepMind Control Suite benchmark for in-distribution performance and CausalWorld benchmark for out-of-distribution generalization. Our code is available at https://github.com/ZexuSun/OILCA-NeurIPS23.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Deep Structural Knowledge Exploitation and Synergy for Estimating Node Importance Value on Heterogeneous Information NetworksYankai Chen, Yixiang Fang, Qiongyan Wang, Xin Cao 等AAAI 2024 · 被引用 17 次
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng 等NeurIPS 2025 · 被引用 7 次
- DIDI: Diffusion-Guided Diversity for Offline Behavioral GenerationJinxin Liu, Xinghong Guo, Zifeng Zhuang, Donglin WangICML 2024 · 被引用 3 次
它引用的顶会 Paper10
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 被引用 105 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
- Explaining the Efficacy of Counterfactually Augmented DataDivyansh Kaushik, Amrith Setlur, Eduard H. Hovy, Zachary Chase LiptonICLR 2021 · 被引用 89 次
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
相关 Paper
- Causal Action Influence Aware Counterfactual Data AugmentationNúria Armengol Urpí, Marco Bagatella, Marin Vlastelica, Georg MartiusICML 2024 · 被引用 11 次
- Offline Imitation Learning with Model-based Reverse AugmentationJie-Jing Shao, Hao-Sen Shi, Lan-Zhe Guo, Yu-Feng LiKDD 2024 · 被引用 5 次
- Mitigating Covariate Shift in Imitation Learning via Offline Data With Partial CoverageJonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi 等NeurIPS 2021 · 被引用 90 次
- Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data CorruptionsRui Yang, Jie Wang, Guoping Wu, Bin LiNeurIPS 2024 · 被引用 11 次
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon 等ICLR 2022 · 被引用 111 次
