Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
Yuping Luo, Huazhe Xu, Tengyu Ma
Abstract
Imitation learning, followed by reinforcement learning algorithms, is a promising paradigm to solve complex control tasks sample-efficiently. However, learning from demonstrations often suffers from the covariate shift problem, which results in cascading errors of the learned policy. We introduce a notion of conservatively-extrapolated value functions, which provably lead to policies with self-correction. We design an algorithm Value Iteration with Negative Sampling (VINS) that practically learns such value functions with conservative extrapolation. We show that VINS can correct mistakes of the behavioral cloning policy on simulated robotics benchmark tasks. We also propose the algorithm of using VINS to initialize a reinforcement learning algorithm, which is shown to outperform significantly prior works in sample efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b182bbf-1a84-4b70-ad00-5f800be74a5bCited by top-tier papers5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 137 citations
- Social NCE: Contrastive Learning of Socially-aware Motion RepresentationsYuejiang Liu, Qi Yan, Alexandre AlahiICCV 2021 · 118 citations
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu et al.NeurIPS 2021 · 28 citations
- Robust Imitation of a Few Demonstrations with a Backwards ModelJung Yeon Park, Lawson L. S. WongNeurIPS 2022 · 21 citations
Related papers
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 112 citations
- CLARE: Conservative Model-Based Reward Learning for Offline Inverse Reinforcement LearningSheng Yue, Guanbo Wang, Wei Shao, Zhaofeng Zhang et al.ICLR 2023 · 6 citations
- Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouAAAI 2022 · 14 citations
- Planning for Sample Efficient Imitation LearningZhao-Heng Yin, Weirui Ye, Qifeng Chen, Yang GaoNeurIPS 2022 · 32 citations
- Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionSeokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang et al.NeurIPS 2024 · 17 citations
