Imitation Learning by Reinforcement Learning
Kamil Ciosek
2022年份
22被引次数
6顶会引用
摘要
Imitation learning algorithms learn a policy from demonstrations of expert behavior. We show that, for deterministic experts, imitation learning can be done by reduction to reinforcement learning with a stationary reward. Our theoretical analysis both certifies the recovery of expert reward and bounds the total variation distance between the expert and the imitation learner, showing a link to adversarial imitation learning. We conduct experiments which confirm that our reduction works well in practice for continuous control tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Fast Imitation via Behavior Foundation ModelsMatteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric 等ICLR 2024 · 被引用 26 次
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 等ICLR 2024 · 被引用 17 次
- Adversarial Moment-Matching Distillation of Large Language ModelsChen JiaNeurIPS 2024 · 被引用 4 次
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie 等ICLR 2026 · 被引用 3 次
- Deep Demonstration Tracing: Learning Generalizable Imitator Policy for Runtime Imitation from a Single DemonstrationXiong-Hui Chen, Junyin Ye, Hang Zhao, Yi-Chen Li 等ICML 2024 · 被引用 2 次
相关 Paper
- Primal Wasserstein Imitation LearningRobert Dadashi, Léonard Hussenot, Matthieu Geist, Olivier PietquinICLR 2021 · 被引用 41 次
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 被引用 112 次
- Imitation with Neural Density ModelsKuno Kim, Akshat Jindal, Yang Song, Jiaming Song 等NeurIPS 2021 · 被引用 14 次
- Variational Adversarial Kernel Learned Imitation LearningFan Yang, Alina Vereshchaka, Yufan Zhou, Changyou Chen 等AAAI 2020 · 被引用 9 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
