Sequential Causal Imitation Learning with Unobserved Confounders
Daniel Kumor, Junzhe Zhang, Elias Bareinboim
摘要
"Monkey see monkey do"is an age-old adage, referring to naïve imitation without a deep understanding of a system's underlying mechanics. Indeed, if a demonstrator has access to information unavailable to the imitator (monkey), such as a different set of sensors, then no matter how perfectly the imitator models its perceived environment (See), attempting to reproduce the demonstrator's behavior (Do) can lead to poor outcomes. Imitation learning in the presence of a mismatch between demonstrator and imitator has been studied in the literature under the rubric of causal imitation learning (Zhang et al., 2020), but existing solutions are limited to single-stage decision-making. This paper investigates the problem of causal imitation learning in sequential settings, where the imitator must make multiple decisions per episode. We develop a graphical criterion that is necessary and sufficient for determining the feasibility of causal imitation, providing conditions when an imitator can match a demonstrator's performance despite differing capabilities. Finally, we provide an efficient algorithm for determining imitability and corroborate our theory with simulations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Sequence Model Imitation Learning with Unobserved ContextsGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven WuNeurIPS 2022 · 被引用 39 次
- Causal Imitation Learning under Temporally Correlated NoiseGokul Swamy, Sanjiban Choudhury, Drew Bagnell, Steven WuICML 2022 · 被引用 36 次
- Adaptively Exploiting d-Separators with Causal BanditsBlair L. Bilodeau, Linbo Wang, Daniel M. RoyNeurIPS 2022 · 被引用 25 次
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
- A Relational Intervention Approach for Unsupervised Dynamics Generalization in Model-Based Reinforcement LearningJiaxian Guo, Mingming Gong, Dacheng TaoICLR 2022 · 被引用 21 次
它引用的顶会 Paper2
相关 Paper
- Causal Imitation for Markov Decision Processes: a Partial Identification ApproachKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimNeurIPS 2024 · 被引用 12 次
- Causal Imitability Under Context-Specific Independence RelationsFateme Jamshidi, Sina Akbari, Negar KiyavashNeurIPS 2023 · 被引用 8 次
- Domain Adaptive Imitation LearningKuno Kim, Yihong Gu, Jiaming Song, Shengjia Zhao 等ICML 2020 · 被引用 86 次
- Learning Human Driving Behaviors with Sequential Causal Imitation LearningKangrui Ruan, Xuan DiAAAI 2022 · 被引用 28 次
- Invariant Causal Imitation Learning for Generalizable PoliciesIoana Bica, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2021 · 被引用 46 次
