Causal Imitation Learning under Temporally Correlated Noise
Gokul Swamy, Sanjiban Choudhury, Drew Bagnell, Steven Wu
摘要
We develop algorithms for imitation learning from policy data that was corrupted by temporally correlated noise in expert actions. When noise affects multiple timesteps of recorded data, it can manifest as spurious correlations between states and actions that a learner might latch on to, leading to poor policy performance. To break up these spurious correlations, we apply modern variants of the instrumental variable regression (IVR) technique of econometrics, enabling us to recover the underlying policy without requiring access to an interactive expert. In particular, we present two techniques, one of a generative-modeling flavor (DoubIL) that can utilize access to a simulator, and one of a game-theoretic flavor (ResiduIL) that can be run entirely offline. We find both of our algorithms compare favorably to behavioral cloning on simulated control tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- QueST: Self-Supervised Skill Abstractions for Learning Continuous ControlAtharva Mete, Haotian Xue, Albert Wilcox, Yongxin Chen 等NeurIPS 2024 · 被引用 76 次
- Sequence Model Imitation Learning with Unobserved ContextsGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven WuNeurIPS 2022 · 被引用 39 次
- Multi-Agent Imitation Learning: Value is Easy, Regret is HardJingwu Tang, Gokul Swamy, Fei Fang, Zhiwei Steven WuNeurIPS 2024 · 被引用 13 次
- Causal Imitation for Markov Decision Processes: a Partial Identification ApproachKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimNeurIPS 2024 · 被引用 12 次
- Causal Imitability Under Context-Specific Independence RelationsFateme Jamshidi, Sina Akbari, Negar KiyavashNeurIPS 2023 · 被引用 8 次
它引用的顶会 Paper5
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 被引用 666 次
- Minimax Estimation of Conditional Moment ModelsNishanth Dikkala, Greg Lewis, Lester Mackey, Vasilis SyrgkanisNeurIPS 2020 · 被引用 125 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
- Causal Imitation Learning With Unobserved ConfoundersJunzhe Zhang, Daniel Kumor, Elias BareinboimNeurIPS 2020 · 被引用 86 次
- Sequential Causal Imitation Learning with Unobserved ConfoundersDaniel Kumor, Junzhe Zhang, Elias BareinboimNeurIPS 2021 · 被引用 53 次
相关 Paper
- Causal Imitation Learning under Expert-Observable and Expert-Unobservable ConfoundingDaqian Shao, Thomas Kleine Buening, Marta KwiatkowskaICLR 2026 · 被引用 1 次
- Learning Decision Policies with Instrumental Variables through Double Machine LearningDaqian Shao, Ashkan Soleymani, Francesco Quinzan, Marta KwiatkowskaICML 2024 · 被引用 4 次
- Fighting Copycat Agents in Behavioral Cloning from Observation HistoriesChuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman 等NeurIPS 2020 · 被引用 103 次
- Learning Human Driving Behaviors with Sequential Causal Imitation LearningKangrui Ruan, Xuan DiAAAI 2022 · 被引用 28 次
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo 等ICML 2023 · 被引用 24 次
