Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPs
Antoine Moulin, Gergely Neu, Luca Viano
摘要
We study the problem of offline imitation learning in Markov decision processes (MDPs), where the goal is to learn a well-performing policy given a dataset of state-action pairs generated by an expert policy. Complementing a recent line of work on this topic that assumes the expert belongs to a tractable class of known policies, we approach this problem from a new angle and leverage a different type of structural assumption about the environment. Specifically, for the class of linear -realizable MDPs, we introduce a new algorithm called saddle-point offline imitation learning (), which is guaranteed to match the performance of any expert up to an additive error with access to samples. Moreover, we extend this result to possibly nonlinear -realizable MDPs at the cost of a worse sample complexity of order . Finally, our analysis suggests a new loss function for training critic networks from expert data in deep imitation learning. Empirical evaluations on standard benchmarks demonstrate that the neural net implementation of is superior to behavior cloning and competitive with state-of-the-art algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning to Answer from Correct DemonstrationsNirmit Joshi, Gene Li, Siddharth Bhandari, Shiva Prasad Kasiviswanathan 等ICLR 2026 · 被引用 6 次
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 被引用 2 次
- Multi-agent imitation learning with function approximation: linear Markov games and beyondLuca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher 等ICML 2026 · 被引用 1 次
- Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow MechanismTian Xu, Chenyang Wang, Xiaochen Zhai, Ziniu Li 等ICML 2026
- Reward Model Evaluation via Automatically-Ranked Policy AlignmentAoran Wang, Lei Ou, Yang Yu, Zongzhang ZhangAAAI 2026
它引用的顶会 Paper14
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 被引用 137 次
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
- Provable Guarantees for Generative Behavior Cloning: Bridging Low-Level Stability and High-Level BehaviorAdam Block, Ali Jadbabaie, Daniel Pfrommer, Max Simchowitz 等NeurIPS 2023 · 被引用 44 次
- Online Apprenticeship LearningLior Shani, Tom Zahavy, Shie MannorAAAI 2022 · 被引用 33 次
相关 Paper
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu 等NeurIPS 2021 · 被引用 28 次
- Mitigating Covariate Shift in Imitation Learning via Offline Data With Partial CoverageJonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi 等NeurIPS 2021 · 被引用 90 次
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 被引用 105 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 被引用 84 次
