RLIF: Interactive Imitation Learning as Reinforcement Learning
Jianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma, Sergey Levine
摘要
Although reinforcement learning methods offer a powerful framework for automatic skill acquisition, for practical learning-based control problems in domains such as robotics, imitation learning often provides a more convenient and accessible alternative. In particular, an interactive imitation learning method such as DAgger, which queries a near-optimal expert to intervene online to collect correction data for addressing the distributional shift challenges that afflict naïve behavioral cloning, can enjoy good performance both in theory and practice without requiring manually specified reward functions and other components of full reinforcement learning methods. In this paper, we explore how off-policy reinforcement learning can enable improved performance under assumptions that are similar but potentially even more practical than those of interactive imitation learning. Our proposed method uses reinforcement learning with user intervention signals themselves as rewards. This relaxes the assumption that intervening experts in interactive imitation learning should be near-optimal and enables the algorithm to learn behaviors that improve over the potential suboptimal human expert. We also provide a unified framework to analyze our RL method and DAgger; for which we present the asymptotic analysis of the suboptimal gap for both methods as well as the nonasymptotic sample complexity bound of our method. We then evaluate our method on challenging high-dimensional continuous control simulation benchmarks as well as real-world robotic vision-based manipulation tasks. The results show that it strongly outperforms DAgger-like approaches across the different tasks, especially when the intervening experts are suboptimal. Additional ablations also empirically verify the proposed theoretical justification that the performance of our method is associated with the choice of intervention model and suboptimality of the expert. Code and videos can be found on the project website: rlif-page.github.io * Equal contributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Human-assisted Robotic Policy Refinement via Action Preference OptimizationWenke Xia, Yichu Yang, Hongtao Wu, Xiao Ma 等NeurIPS 2025 · 被引用 17 次
- Progressive Learning with Human Feedback for Personalized Adaptive Video StreamingZhaohui Jiang, Xuening Feng, Tianchi Huang, Ruixiao Zhang 等ACM MM 2025
- Policy Optimization under Imperfect Human Interactions with Agent-Gated Shared AutonomyZhenghai Xue, Bo An, Shuicheng YanICLR 2025
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu 等ICLR 2025
- Discriminator-Guided Embodied Planning for LLM AgentHaofu Qian, Chenjia Bai, Jiatao Zhang, Fei Wu 等ICLR 2025
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
相关 Paper
- Interactive and Hybrid Imitation Learning: Provably Beating Behavior CloningYichen Li, Chicheng ZhangNeurIPS 2025 · 被引用 1 次
- TaSIL: Taylor Series Imitation LearningDaniel Pfrommer, Thomas T. C. K. Zhang, Stephen Tu, Nikolai MatniNeurIPS 2022 · 被引用 27 次
- Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human CorrectionsXiaomeng Xu, Yifan Hou, Zeyi Liu, Shuran SongNeurIPS 2025 · 被引用 57 次
- Agnostic Interactive Imitation Learning: New Theory and Practical AlgorithmsYichen Li, Chicheng ZhangICML 2024
- Fine-tuning Behavioral Cloning Policies with Preference‑Based Reinforcement LearningMaël Macuglia, Paul Friedrich, Giorgia RamponiICLR 2026 · 被引用 2 次
