Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation Gap
Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven Wu
摘要
We provide a unifying view of a large family of previous imitation learning algorithms through the lens of moment matching. At its core, our classification scheme is based on whether the learner attempts to match (1) reward or (2) action-value moments of the expert's behavior, with each option leading to differing algorithmic approaches. By considering adversarially chosen divergences between learner and expert behavior, we are able to derive bounds on policy performance that apply for all algorithms in each of these classes, the first to our knowledge. We also introduce the notion of moment recoverability, implicit in many previous analyses of imitation learning, which allows us to cleanly delineate how well each algorithmic family is able to mitigate compounding errors. We derive three novel algorithm templates (AdVIL, AdRIL, and DAeQuIL) with strong guarantees, simple implementation, and competitive empirical performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper53
- Adversarially Trained Actor Critic for Offline Reinforcement LearningChing-An Cheng, Tengyang Xie, Nan Jiang, Alekh AgarwalICML 2022 · 被引用 156 次
- A Minimaximalist Approach to Reinforcement Learning from Human FeedbackGokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu 等ICML 2024 · 被引用 147 次
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- REBEL: Reinforcement Learning via Regressing Relative RewardsZhaolin Gao, Jonathan D. Chang, Wenhao Zhan, Owen Oertell 等NeurIPS 2024 · 被引用 82 次
- All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-TuningGokul Swamy, Sanjiban Choudhury, Wen Sun, Steven Wu 等ICLR 2026 · 被引用 66 次
它引用的顶会 Paper8
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 被引用 137 次
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 被引用 112 次
- Strictly Batch Imitation Learning by Energy-based Distribution MatchingDaniel Jarrett, Ioana Bica, Mihaela van der SchaarNeurIPS 2020 · 被引用 74 次
相关 Paper
- Primal Wasserstein Imitation LearningRobert Dadashi, Léonard Hussenot, Matthieu Geist, Olivier PietquinICLR 2021 · 被引用 41 次
- Imitation Learning by Reinforcement LearningKamil CiosekICLR 2022 · 被引用 22 次
- Causal Imitation Learning under Temporally Correlated NoiseGokul Swamy, Sanjiban Choudhury, Drew Bagnell, Steven WuICML 2022 · 被引用 36 次
- Policy Contrastive Imitation LearningJialei Huang, Zhao-Heng Yin, Yingdong Hu, Yang GaoICML 2023 · 被引用 4 次
- Causal Imitation Learning via Inverse Reinforcement LearningKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimICLR 2023
