Lune

ICLR2020Top-tier venue

Disagreement-Regularized Imitation Learning

Kianté Brantley, Wen Sun, Mikael Henaff

2020Year
112Citations
43Top-tier citations

Abstract

We present a simple and effective algorithm designed to address the covariate shift problem in imitation learning. It operates by training an ensemble of policies on the expert demonstration data, and using the variance of their predictions as a cost which is minimized with RL together with a supervised behavioral cloning cost. Unlike adversarial imitation methods, it uses a fixed reward function which is easy to optimize. We prove a regret bound for the algorithm which is linear in the time horizon multiplied by a coefficient which we show to be low for certain problems in which behavioral cloning fails. We evaluate our algorithm empirically across multiple pixel-based Atari environments and continuous control tasks, and show that it matches or significantly outperforms behavioral cloning and generative adversarial imitation learning.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 37df3ba0-9b25-44c1-8e9e-13ef294c8d4e

Cited by top-tier papers43

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines