Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
Yuping Luo, Huazhe Xu, Tengyu Ma
摘要
Imitation learning, followed by reinforcement learning algorithms, is a promising paradigm to solve complex control tasks sample-efficiently. However, learning from demonstrations often suffers from the covariate shift problem, which results in cascading errors of the learned policy. We introduce a notion of conservatively-extrapolated value functions, which provably lead to policies with self-correction. We design an algorithm Value Iteration with Negative Sampling (VINS) that practically learns such value functions with conservative extrapolation. We show that VINS can correct mistakes of the behavioral cloning policy on simulated robotics benchmark tasks. We also propose the algorithm of using VINS to initialize a reinforcement learning algorithm, which is shown to outperform significantly prior works in sample efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 被引用 137 次
- Social NCE: Contrastive Learning of Socially-aware Motion RepresentationsYuejiang Liu, Qi Yan, Alexandre AlahiICCV 2021 · 被引用 118 次
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu 等NeurIPS 2021 · 被引用 28 次
- Robust Imitation of a Few Demonstrations with a Backwards ModelJung Yeon Park, Lawson L. S. WongNeurIPS 2022 · 被引用 21 次
相关 Paper
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 被引用 112 次
- CLARE: Conservative Model-Based Reward Learning for Offline Inverse Reinforcement LearningSheng Yue, Guanbo Wang, Wei Shao, Zhaofeng Zhang 等ICLR 2023 · 被引用 6 次
- Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouAAAI 2022 · 被引用 14 次
- Planning for Sample Efficient Imitation LearningZhao-Heng Yin, Weirui Ye, Qifeng Chen, Yang GaoNeurIPS 2022 · 被引用 32 次
- Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionSeokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang 等NeurIPS 2024 · 被引用 17 次
