Sequential Causal Imitation Learning with Unobserved Confounders
Daniel Kumor, Junzhe Zhang, Elias Bareinboim
Abstract
"Monkey see monkey do"is an age-old adage, referring to naïve imitation without a deep understanding of a system's underlying mechanics. Indeed, if a demonstrator has access to information unavailable to the imitator (monkey), such as a different set of sensors, then no matter how perfectly the imitator models its perceived environment (See), attempting to reproduce the demonstrator's behavior (Do) can lead to poor outcomes. Imitation learning in the presence of a mismatch between demonstrator and imitator has been studied in the literature under the rubric of causal imitation learning (Zhang et al., 2020), but existing solutions are limited to single-stage decision-making. This paper investigates the problem of causal imitation learning in sequential settings, where the imitator must make multiple decisions per episode. We develop a graphical criterion that is necessary and sufficient for determining the feasibility of causal imitation, providing conditions when an imitator can match a demonstrator's performance despite differing capabilities. Finally, we provide an efficient algorithm for determining imitability and corroborate our theory with simulations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de46a6fa-e0ad-4fb5-a5e1-71ed2fec5450Cited by top-tier papers14
- Sequence Model Imitation Learning with Unobserved ContextsGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven WuNeurIPS 2022 · 39 citations
- Causal Imitation Learning under Temporally Correlated NoiseGokul Swamy, Sanjiban Choudhury, Drew Bagnell, Steven WuICML 2022 · 36 citations
- Adaptively Exploiting d-Separators with Causal BanditsBlair L. Bilodeau, Linbo Wang, Daniel M. RoyNeurIPS 2022 · 25 citations
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 22 citations
- A Relational Intervention Approach for Unsupervised Dynamics Generalization in Model-Based Reinforcement LearningJiaxian Guo, Mingming Gong, Dacheng TaoICLR 2022 · 21 citations
Builds on2
Related papers
- Causal Imitation for Markov Decision Processes: a Partial Identification ApproachKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimNeurIPS 2024 · 12 citations
- Causal Imitability Under Context-Specific Independence RelationsFateme Jamshidi, Sina Akbari, Negar KiyavashNeurIPS 2023 · 8 citations
- Domain Adaptive Imitation LearningKuno Kim, Yihong Gu, Jiaming Song, Shengjia Zhao et al.ICML 2020 · 86 citations
- Learning Human Driving Behaviors with Sequential Causal Imitation LearningKangrui Ruan, Xuan DiAAAI 2022 · 28 citations
- Invariant Causal Imitation Learning for Generalizable PoliciesIoana Bica, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2021 · 46 citations
