Lune

NeurIPS2020Top-tier venue

Task-agnostic Exploration in Reinforcement Learning

Xuezhou Zhang, Yuzhe Ma, Adish Singla

2020Year
56Citations
27Top-tier citations

Abstract

Efficient exploration is one of the main challenges in reinforcement learning (RL). Most existing sample-efficient algorithms assume the existence of a single reward function during exploration. In many practical scenarios, however, there is not a single underlying reward function to guide the exploration, for instance, when an agent needs to learn many skills simultaneously, or multiple conflicting objectives need to be balanced. To address these challenges, we propose the task-agnostic RL framework: In the exploration phase, the agent first collects trajectories by exploring the MDP without the guidance of a reward function. After exploration, it aims at finding near-optimal policies for NN tasks, given the collected trajectories augmented with sampled rewards for each task. We present an efficient task-agnostic RL algorithm, UCBZero, that finds ϵ\epsilon-optimal policies for NN arbitrary tasks after at most O~(log⁡(N)H5SA/ϵ2)\tilde O(\log(N)H^5SA/\epsilon^2) exploration episodes. We also provide an Ω(log⁡(N)H2SA/ϵ2)\Omega(\log (N)H^2SA/\epsilon^2) lower bound, showing that the log⁡\log dependency on NN is unavoidable. Furthermore, we provide an NN-independent sample complexity bound of UCBZero in the statistically easier setting when the ground truth reward functions are known.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 686df2ee-2024-41f7-bec1-ba7ebb1db479

Cited by top-tier papers27

Ask how each one uses it

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines