Lune

ICLR2020Top-tier venue

Sub-policy Adaptation for Hierarchical Reinforcement Learning

Alexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter Abbeel

2020Year
85Citations
20Top-tier citations

Abstract

Hierarchical reinforcement learning is a promising approach to tackle long-horizon decision-making problems with sparse rewards. Unfortunately, most methods still decouple the lower-level skill acquisition process and the training of a higher level that controls the skills in a new task. Leaving the skills fixed can lead to significant sub-optimality in the transfer setting. In this work, we propose a novel algorithm to discover a set of skills, and continuously adapt them along with the higher level even when training on a new task. Our main contributions are two-fold. First, we derive a new hierarchical policy gradient with an unbiased latent-dependent baseline, and we introduce Hierarchical Proximal Policy Optimization (HiPPO), an on-policy method to efficiently train all levels of the hierarchy jointly. Second, we propose a method for training time-abstractions that improves the robustness of the obtained skills to environment changes. Code and results are available at this http URL

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext fee0d45f-7d64-4081-8535-6d91e234be47

Cited by top-tier papers20

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines