Lune

NeurIPS2025Top-tier venue

Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning

Yichen Li, Chicheng Zhang

2025Year
1Citations

Abstract

Imitation learning (IL) is a paradigm for learning sequential decision-making policies from experts, leveraging offline demonstrations, interactive annotations, or both. Recent advances show that when annotation cost is tallied per trajectory, Behavior Cloning (BC)-which relies solely on offline demonstrations-cannot be improved in general, leaving limited conditions for interactive methods such as DAgger to help. We revisit this conclusion and prove that when the annotation cost is measured per state, algorithms using interactive annotations can provably outperform BC. Specifically: (1) we show that STAGGER, a one-sample-per-round variant of DAgger, provably beats BC under low-recovery-cost settings; (2) we initiate the study of hybrid IL where the agent learns from offline demonstrations and interactive annotations. We propose WARM-STAGGER whose learning guarantee is not much worse than using either data source alone. Furthermore, motivated by compounding error and cold-start problem in imitation learning practice, we give an MDP example in which WARM-STAGGER has significant better annotation cost; (3) experiments on MuJoCo continuous-control tasks confirm that, with modest cost ratio between interactive and offline annotations, interactive and hybrid approaches consistently outperform BC. To the best of our knowledge, our work is the first to highlight the benefit of state-wise interactive annotation and hybrid feedback in imitation learning.

Recent work [17], via a sharp analysis of Behavior Cloning, shows that the sample efficiency of Behavior Cloning cannot be improved in general when measuring using the number of trajectories annotated. Interactive methods like DAgger [49] can enjoy sample complexity benefits, but so far the benefits are only exhibited in limited examples, with the most general ones in the tabular setting [46]. This leaves open the question: Can interaction provide sample efficiency benefit for imitation learning under diverse settings, notably in the presence of function approximation?

In this paper, we make progress towards this question, with a focus on the deterministically realizable setting (i.e. the expert policy π E is deterministic and is in the learner's policy class B). Specifically, we make the following contributions:

  1. Motivated by the costly nature of interactive labeling on entire trajectories [29,37], we propose to measure the cost of annotation using the number of states annotated by the demonstrating expert. We propose a general state-wise interactive imitation learning algorithm, STAGGER, and show that as long as the expert can recover from mistakes at low cost in the environment [51], it significantly improves over Behavior Cloning in terms of its number of state-wise demonstrations required. 2. Motivated by practical imitation learning applications where sets of offline demonstration data are readily available, we study hybrid imitation learning, where the learning agent can additionally query the demonstration expert interactively to improve its performance. We design a hybrid imitation learning algorithm, WARM-STAGGER, and prove that its policy optimality guarantee is not much worse than using either of the data sources alone. 3. Inspired by compounding error [45] and cold start problem [35,42], we provide an MDP example, for which we show hybrid imitation learning can achieve strict sample complexity savings over using either source alone, and provide simulations that verify this claim. 4. We conduct experiments in MuJoCo continuous control tasks and show that if the cost of state-wise interactive demonstration is not much higher than its offline counterpart, interactive algorithms can enjoy a better cost efficiency than Behavior Cloning. Under some cost regimes and some environments, hybrid imitation learning can outperform approaches that use either source alone.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f74702b7-d2d4-45c3-aae3-42d4a9f9f43f

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines