Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
Yichen Li, Chicheng Zhang
Abstract
Imitation learning (IL) is a paradigm for learning sequential decision-making policies from experts, leveraging offline demonstrations, interactive annotations, or both. Recent advances show that when annotation cost is tallied per trajectory, Behavior Cloning (BC)-which relies solely on offline demonstrations-cannot be improved in general, leaving limited conditions for interactive methods such as DAgger to help. We revisit this conclusion and prove that when the annotation cost is measured per state, algorithms using interactive annotations can provably outperform BC. Specifically: (1) we show that STAGGER, a one-sample-per-round variant of DAgger, provably beats BC under low-recovery-cost settings; (2) we initiate the study of hybrid IL where the agent learns from offline demonstrations and interactive annotations. We propose WARM-STAGGER whose learning guarantee is not much worse than using either data source alone. Furthermore, motivated by compounding error and cold-start problem in imitation learning practice, we give an MDP example in which WARM-STAGGER has significant better annotation cost; (3) experiments on MuJoCo continuous-control tasks confirm that, with modest cost ratio between interactive and offline annotations, interactive and hybrid approaches consistently outperform BC. To the best of our knowledge, our work is the first to highlight the benefit of state-wise interactive annotation and hybrid feedback in imitation learning.
Recent work [17], via a sharp analysis of Behavior Cloning, shows that the sample efficiency of Behavior Cloning cannot be improved in general when measuring using the number of trajectories annotated. Interactive methods like DAgger [49] can enjoy sample complexity benefits, but so far the benefits are only exhibited in limited examples, with the most general ones in the tabular setting [46]. This leaves open the question: Can interaction provide sample efficiency benefit for imitation learning under diverse settings, notably in the presence of function approximation?
In this paper, we make progress towards this question, with a focus on the deterministically realizable setting (i.e. the expert policy π E is deterministic and is in the learner's policy class B). Specifically, we make the following contributions:
- Motivated by the costly nature of interactive labeling on entire trajectories [29,37], we propose to measure the cost of annotation using the number of states annotated by the demonstrating expert. We propose a general state-wise interactive imitation learning algorithm, STAGGER, and show that as long as the expert can recover from mistakes at low cost in the environment [51], it significantly improves over Behavior Cloning in terms of its number of state-wise demonstrations required. 2. Motivated by practical imitation learning applications where sets of offline demonstration data are readily available, we study hybrid imitation learning, where the learning agent can additionally query the demonstration expert interactively to improve its performance. We design a hybrid imitation learning algorithm, WARM-STAGGER, and prove that its policy optimality guarantee is not much worse than using either of the data sources alone. 3. Inspired by compounding error [45] and cold start problem [35,42], we provide an MDP example, for which we show hybrid imitation learning can achieve strict sample complexity savings over using either source alone, and provide simulations that verify this claim. 4. We conduct experiments in MuJoCo continuous control tasks and show that if the cost of state-wise interactive demonstration is not much higher than its offline counterpart, interactive algorithms can enjoy a better cost efficiency than Behavior Cloning. Under some cost regimes and some environments, hybrid imitation learning can outperform approaches that use either source alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f74702b7-d2d4-45c3-aae3-42d4a9f9f43fBuilds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 137 citations
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 112 citations
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 112 citations
- Mitigating Covariate Shift in Imitation Learning via Offline Data With Partial CoverageJonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi et al.NeurIPS 2021 · 90 citations
Related papers
- Agnostic Interactive Imitation Learning: New Theory and Practical AlgorithmsYichen Li, Chicheng ZhangICML 2024
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
- TaSIL: Taylor Series Imitation LearningDaniel Pfrommer, Thomas T. C. K. Zhang, Stephen Tu, Nikolai MatniNeurIPS 2022 · 27 citations
- On Efficient Online Imitation Learning via ClassificationYichen Li, Chicheng ZhangNeurIPS 2022 · 7 citations
- Robot-Gated Interactive Imitation Learning with Adaptive Intervention MechanismHaoyuan Cai, Zhenghao Peng, Bolei ZhouICML 2025
