Toward the Fundamental Limits of Imitation Learning
Nived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan Ramchandran
Abstract
Imitation learning (IL) aims to mimic the behavior of an expert policy in a sequential decision-making problem given only demonstrations. In this paper, we focus on understanding the minimax statistical limits of IL in episodic Markov Decision Processes (MDPs). We first consider the setting where the learner is provided a dataset of expert trajectories ahead of time, and cannot interact with the MDP. Here, we show that the policy which mimics the expert whenever possible is in expectation suboptimal compared to the value of the expert, even when the expert follows an arbitrary stochastic policy. Here is the state space, and is the length of the episode. Furthermore, we establish a suboptimality lower bound of which applies even if the expert is constrained to be deterministic, or if the learner is allowed to actively query the expert at visited states while interacting with the MDP for episodes. To our knowledge, this is the first algorithm with suboptimality having no dependence on the number of actions, under no additional assumptions. We then propose a novel algorithm based on minimum-distance functionals in the setting where the transition model is given and the expert is deterministic. The algorithm is suboptimal by , showing that knowledge of transition improves the minimax rate by at least a factor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 399ec115-e6ce-466d-9292-ebc6bcef6d58Cited by top-tier papers60
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 176 citations
- Behavior Generation with Latent ActionsSeungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim et al.ICML 2024 · 154 citations
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 112 citations
- Pessimistic Q-Learning for Offline Reinforcement Learning: Towards Optimal Sample ComplexityLaixi Shi, Gen Li, Yuting Wei, Yuxin Chen et al.ICML 2022 · 110 citations
Builds on3
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 112 citations
- Provable Representation Learning for Imitation Learning via Bi-level OptimizationSanjeev Arora, Simon S. Du, Sham M. Kakade, Yuping Luo et al.ICML 2020 · 65 citations
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative SamplingYuping Luo, Huazhe Xu, Tengyu MaICLR 2020 · 14 citations
Related papers
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu et al.NeurIPS 2021 · 28 citations
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 2 citations
- Imitation Learning in Discounted Linear MDPs without exploration assumptionsLuca Viano, Stratis Skoulakis, Volkan CevherICML 2024 · 10 citations
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 56 citations
- Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPsAntoine Moulin, Gergely Neu, Luca VianoNeurIPS 2025 · 6 citations
