On the Value of Interaction and Function Approximation in Imitation Learning
Nived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu, Jiantao Jiao, Kannan Ramchandran
摘要
We study the statistical guarantees for the Imitation Learning (IL) problem in episodic MDPs. [22] show an information theoretic lower bound that in the worst case, a learner which can even actively query the expert policy suffers from a suboptimality growing quadratically in the length of the horizon, H. We study imitation learning under the µ-recoverability assumption of [27] which assumes that the difference in the Q-value under the expert policy across different actions in a state do not deviate beyond µ from the maximum. We show that the reduction proposed by [25] is statistically optimal: the resulting algorithm upon interacting with the MDP for N episodes results in a suboptimality bound of O (µ|S|H/N ) which we show is optimal up to log-factors. In contrast, we show that any algorithm which does not interact with the MDP and uses an offline dataset of N expert trajectories must incur suboptimality growing as |S|H 2 /N even under the µrecoverability assumption. This establishes a clear and provable separation of the minimax rates between the active setting and the no-interaction setting. We also study IL with linear function approximation. When the expert plays actions according to a linear classifier of known state-action features, we use the reduction to multi-class classification to show that with high probability, the suboptimality of behavior cloning is O(dH 2 /N ) given N rollouts from the optimal policy. This is optimal up to log-factors but can be improved to O(dH/N ) if we have a linear expert with parameter-sharing across time steps. In contrast, when the MDP transition structure is known to the learner such as in the case of simulators, we demonstrate fundamental differences compared to the tabular setting in terms of the performance of an optimal algorithm, MIMIC-MD (Rajaraman et al. [22]) when extended to the function approximation setting. Here, we introduce a new problem called confidence set linear classification, that can be used to construct sample-efficient IL algorithms. H t =t r t (s t , a t )|s t = s, a t = a], and f π t (s) is defined as the distribution over states induced at time t, by rolling out the policy π.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- Minimax Optimal Online Imitation Learning via Replay EstimationGokul Swamy, Nived Rajaraman, Matthew Peng, Sanjiban Choudhury 等NeurIPS 2022 · 被引用 27 次
- Imitation Learning from Imperfection: Theoretical Justifications and AlgorithmsZiniu Li, Tian Xu, Zeyu Qin, Yang Yu 等NeurIPS 2023 · 被引用 26 次
- Online Learning in Stackelberg Games with an Omniscient FollowerGeng Zhao, Banghua Zhu, Jiantao Jiao, Michael I. JordanICML 2023 · 被引用 23 次
- Memory-Consistent Neural Networks for Imitation LearningKaustubh Sridhar, Souradeep Dutta, Dinesh Jayaraman, James Weimer 等ICLR 2024 · 被引用 15 次
它引用的顶会 Paper4
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 被引用 137 次
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 被引用 112 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative SamplingYuping Luo, Huazhe Xu, Tengyu MaICLR 2020 · 被引用 14 次
相关 Paper
- Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPsAntoine Moulin, Gergely Neu, Luca VianoNeurIPS 2025 · 被引用 6 次
- On Efficient Online Imitation Learning via ClassificationYichen Li, Chicheng ZhangNeurIPS 2022 · 被引用 7 次
- Selective Sampling and Imitation Learning via Online RegressionAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 被引用 15 次
- Imitation Learning in Discounted Linear MDPs without exploration assumptionsLuca Viano, Stratis Skoulakis, Volkan CevherICML 2024 · 被引用 10 次
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 被引用 2 次
