Lune

NeurIPS2020顶会

Task-agnostic Exploration in Reinforcement Learning

Xuezhou Zhang, Yuzhe Ma, Adish Singla

2020年份
56被引次数
27顶会引用

摘要

Efficient exploration is one of the main challenges in reinforcement learning (RL). Most existing sample-efficient algorithms assume the existence of a single reward function during exploration. In many practical scenarios, however, there is not a single underlying reward function to guide the exploration, for instance, when an agent needs to learn many skills simultaneously, or multiple conflicting objectives need to be balanced. To address these challenges, we propose the task-agnostic RL framework: In the exploration phase, the agent first collects trajectories by exploring the MDP without the guidance of a reward function. After exploration, it aims at finding near-optimal policies for NN tasks, given the collected trajectories augmented with sampled rewards for each task. We present an efficient task-agnostic RL algorithm, UCBZero, that finds ϵ\epsilon-optimal policies for NN arbitrary tasks after at most O~(log⁡(N)H5SA/ϵ2)\tilde O(\log(N)H^5SA/\epsilon^2) exploration episodes. We also provide an Ω(log⁡(N)H2SA/ϵ2)\Omega(\log (N)H^2SA/\epsilon^2) lower bound, showing that the log⁡\log dependency on NN is unavoidable. Furthermore, we provide an NN-independent sample complexity bound of UCBZero in the statistically easier setting when the ground truth reward functions are known.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper27

问问它们各自怎么用它

它引用的顶会 Paper1

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖