Sequence Model Imitation Learning with Unobserved Contexts
Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven Wu
Abstract
We consider imitation learning problems where the learner's ability to mimic the expert increases throughout the course of an episode as more information is revealed. One example of this is when the expert has access to privileged information: while the learner might not be able to accurately reproduce expert behavior early on in an episode, by considering the entire history of states and actions, they might be able to eventually identify the hidden context and act as the expert would. We prove that on-policy imitation learning algorithms (with or without access to a queryable expert) are better equipped to handle these sorts of asymptotically realizable problems than off-policy methods. This is because onpolicy algorithms provably learn to recover from their initially suboptimal actions, while off-policy methods treat their suboptimal past actions as though they came from the expert. This often manifests as a latching behavior: a naive repetition of past actions. We conduct experiments in a toy bandit domain that show that there exist sharp phase transitions of whether off-policy approaches are able to match expert performance asymptotically, in contrast to the uniformly good performance of on-policy approaches. We demonstrate that on several continuous control tasks, on-policy approaches are able to use history to identify the context while off-policy approaches actually perform worse when given access to history. Recent theoretical work has established "no-go" results for successfully imitating an expert that has access to more information [Zhang et al., 2020 , Kumor et al., 2021] . The core of their arguments is that without seeing some feature that influences expert behavior but is not echoed elsewhere in the state, the learner might not be able to properly ground the expert actions in the observed state. In causal inference terms, this hidden information acts as an unobserved confounder which prevents identification of the desired causal estimand (the expert action). Despite these results, 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d1e1602-5b60-4847-bc99-59ead95c7b1aCited by top-tier papers16
- Self-Distillation Enables Continual LearningIdan Shenfeld, Mehul Damani, Jonas Hübotter, Pulkit AgrawalICML 2026 · 159 citations
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj et al.NeurIPS 2025 · 27 citations
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 22 citations
- Privileged Sensing Scaffolds Reinforcement LearningEdward S. Hu, James Springer, Oleh Rybkin, Dinesh JayaramanICLR 2024 · 21 citations
- Regularized Behavior Cloning for Blocking the Leakage of Past Action InformationSeokin Seo, HyeongJoo Hwang, Hongseok Yang, Kee-Eung KimNeurIPS 2023 · 14 citations
Builds on9
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
- Fighting Copycat Agents in Behavioral Cloning from Observation HistoriesChuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman et al.NeurIPS 2020 · 103 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
- Causal Imitation Learning With Unobserved ConfoundersJunzhe Zhang, Daniel Kumor, Elias BareinboimNeurIPS 2020 · 86 citations
- Sequential Causal Imitation Learning with Unobserved ConfoundersDaniel Kumor, Junzhe Zhang, Elias BareinboimNeurIPS 2021 · 53 citations
Related papers
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
- Causal Imitation for Markov Decision Processes: a Partial Identification ApproachKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimNeurIPS 2024 · 12 citations
- Student-Informed Teacher TrainingNico Messikommer, Jiaxu Xing, Elie Aljalbout, Davide ScaramuzzaICLR 2025
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador et al.NeurIPS 2021 · 53 citations
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu et al.NeurIPS 2021 · 28 citations
