Toward the Fundamental Limits of Imitation Learning
Nived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan Ramchandran
摘要
Imitation learning (IL) aims to mimic the behavior of an expert policy in a sequential decision-making problem given only demonstrations. In this paper, we focus on understanding the minimax statistical limits of IL in episodic Markov Decision Processes (MDPs). We first consider the setting where the learner is provided a dataset of expert trajectories ahead of time, and cannot interact with the MDP. Here, we show that the policy which mimics the expert whenever possible is in expectation suboptimal compared to the value of the expert, even when the expert follows an arbitrary stochastic policy. Here is the state space, and is the length of the episode. Furthermore, we establish a suboptimality lower bound of which applies even if the expert is constrained to be deterministic, or if the learner is allowed to actively query the expert at visited states while interacting with the MDP for episodes. To our knowledge, this is the first algorithm with suboptimality having no dependence on the number of actions, under no additional assumptions. We then propose a novel algorithm based on minimum-distance functionals in the setting where the transition model is given and the expert is deterministic. The algorithm is suboptimal by , showing that knowledge of transition improves the minimax rate by at least a factor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper60
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 被引用 176 次
- Behavior Generation with Latent ActionsSeungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim 等ICML 2024 · 被引用 154 次
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- Pessimistic Q-Learning for Offline Reinforcement Learning: Towards Optimal Sample ComplexityLaixi Shi, Gen Li, Yuting Wei, Yuxin Chen 等ICML 2022 · 被引用 110 次
它引用的顶会 Paper3
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 被引用 112 次
- Provable Representation Learning for Imitation Learning via Bi-level OptimizationSanjeev Arora, Simon S. Du, Sham M. Kakade, Yuping Luo 等ICML 2020 · 被引用 65 次
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative SamplingYuping Luo, Huazhe Xu, Tengyu MaICLR 2020 · 被引用 14 次
相关 Paper
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu 等NeurIPS 2021 · 被引用 28 次
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 被引用 2 次
- Imitation Learning in Discounted Linear MDPs without exploration assumptionsLuca Viano, Stratis Skoulakis, Volkan CevherICML 2024 · 被引用 10 次
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 被引用 56 次
- Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPsAntoine Moulin, Gergely Neu, Luca VianoNeurIPS 2025 · 被引用 6 次
