Imitation Learning in Discounted Linear MDPs without exploration assumptions
Luca Viano, Stratis Skoulakis, Volkan Cevher
Abstract
We present a new algorithm for imitation learning in infinite horizon linear MDPs dubbed ILARL which greatly improves the bound on the number of trajectories that the learner needs to sample from the environment. In particular, we remove exploration assumptions required in previous works and we improve the dependence on the desired accuracy ϵ from O ϵ -5 to O ϵ -4 . Our result relies on a connection between imitation learning and online learning in MDPs with adversarial losses. For the latter setting, we present the first result for infinite horizon linear MDP which may be of independent interest. Moreover, we are able to provide a strengthen result for the finite horizon case where we achieve O ϵ -2 . Numerical experiments with linear function approximation shows that ILARL outperforms other commonly used algorithms. 1 That is, we consider Linear MDP [22] whose formal definition is deferred to Section 2.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e44a0c6-d282-4fe4-8df6-de4e90f5592cCited by top-tier papers9
- Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPsAntoine Moulin, Gergely Neu, Luca VianoNeurIPS 2025 · 6 citations
- How does Inverse RL Scale to Large State Spaces? A Provably Efficient ApproachFilippo Lazzati, Mirco Mutti, Alberto Maria MetelliNeurIPS 2024 · 5 citations
- Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation LearningTill Freihaut, Luca Viano, Volkan Cevher, Matthieu Geist et al.NeurIPS 2025 · 4 citations
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 2 citations
- Imitation Learning as Return Distribution MatchingFilippo Lazzati, Alberto Maria MetelliICLR 2026 · 1 citation
Builds on27
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 271 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Bilinear Classes: A Structural Framework for Provable Generalization in RLSimon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett et al.ICML 2021 · 207 citations
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 161 citations
Related papers
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk et al.NeurIPS 2022 · 27 citations
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu et al.NeurIPS 2021 · 28 citations
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 137 citations
- Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double ExplorationHeyang Zhao, Xingrui Yu, David Mark Bossens, Ivor W. Tsang et al.ICLR 2025
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 112 citations
