Causal Imitation for Markov Decision Processes: a Partial Identification Approach
Kangrui Ruan, Junzhe Zhang, Xuan Di, Elias Bareinboim
Abstract
Imitation learning enables an agent to learn from expert demonstrations when the performance measure is unknown and the reward signal is not specified. Standard imitation methods do not generally apply when the learner and the expert’s sensory capabilities mismatch and demonstrations are contaminated with unobserved confounding bias. To address these challenges, recent advancements in causal imitation learning have been pursued. However, these methods often require access to underlying causal structures that might not always be available, posing practical challenges. In this paper, we investigate robust imitation learning within the framework of canonical Markov Decision Processes (MDPs) using partial identification, allowing the agent to achieve expert performance even when the system dynamics are not uniquely determined from the confounded expert demonstrations. Specifically, first, we theoretically demonstrate that when unobserved confounders (UCs) exist in an MDP, the learner is generally unable to imitate expert performance. We then explore imitation learning in partially identifiable settings — either transition distribution or reward function is non-identifiable from the available data and knowledge. Augmenting the celebrated GAIL method (Ho & Ermon, 2016), our analysis leads to two novel causal imitation algorithms that can obtain effective policies guaranteed to achieve expert performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b0c47bf-2f66-4216-a6cd-c24dd87c5cb1Cited by top-tier papers6
- Confounding Robust Deep Reinforcement Learning: A Causal ApproachMingxuan Li, Junzhe Zhang, Elias BareinboimNeurIPS 2025 · 7 citations
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 1 citation
- Causal Imitation Learning under Expert-Observable and Expert-Unobservable ConfoundingDaqian Shao, Thomas Kleine Buening, Marta KwiatkowskaICLR 2026 · 1 citation
- Polynomial-Delay MAG Listing with Novel Locally Complete Orientation RulesTian-Zuo Wang, Wen-Bo Du, Zhi-Hua ZhouICML 2025
- Counterfactual Bootstrap for Robust Meta-Reinforcement LearningAi Bo, Junzhe Zhang, M. Cenk GursoyICML 2026
Builds on11
- A Calculus for Stochastic Interventions: Causal Effect Identification and Surrogate ExperimentsJuan D. Correa, Elias BareinboimAAAI 2020 · 90 citations
- Causal Imitation Learning With Unobserved ConfoundersJunzhe Zhang, Daniel Kumor, Elias BareinboimNeurIPS 2020 · 86 citations
- Partial Counterfactual Identification from Observational and Experimental DataJunzhe Zhang, Jin Tian, Elias BareinboimICML 2022 · 77 citations
- Sequential Causal Imitation Learning with Unobserved ConfoundersDaniel Kumor, Junzhe Zhang, Elias BareinboimNeurIPS 2021 · 53 citations
- Invariant Causal Imitation Learning for Generalizable PoliciesIoana Bica, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2021 · 46 citations
Related papers
- Causal Imitation Learning via Inverse Reinforcement LearningKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimICLR 2023
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor et al.ICLR 2022 · 16 citations
- Learning Human Driving Behaviors with Sequential Causal Imitation LearningKangrui Ruan, Xuan DiAAAI 2022 · 28 citations
- Fighting Copycat Agents in Behavioral Cloning from Observation HistoriesChuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman et al.NeurIPS 2020 · 103 citations
- Causal Flow Q-Learning for Robust Offline Reinforcement LearningMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2026 · 1 citation
