Lune

ICML2020顶会

Reward-Free Exploration for Reinforcement Learning

Chi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng Yu

2020年份
226被引次数
139顶会引用

摘要

Exploration is widely regarded as one of the most challenging aspects of reinforcement learning (RL), with many naive approaches succumbing to exponential sample complexity. To isolate the challenges of exploration, we propose a new "reward-free RL" framework. In the exploration phase, the agent first collects trajectories from an MDP M\mathcal{M} without a pre-specified reward function. After exploration, it is tasked with computing near-optimal policies under for M\mathcal{M} for a collection of given reward functions. This framework is particularly suitable when there are many reward functions of interest, or when the reward function is shaped by an external agent to elicit desired behavior. We give an efficient algorithm that conducts O~(S2Apoly(H)/ϵ2)\tilde{\mathcal{O}}(S^2A\mathrm{poly}(H)/\epsilon^2) episodes of exploration and returns ϵ\epsilon-suboptimal policies for an arbitrary number of reward functions. We achieve this by finding exploratory policies that visit each "significant" state with probability proportional to its maximum visitation probability under any possible policy. Moreover, our planning procedure can be instantiated by any black-box approximate planner, such as value iteration or natural policy gradient. We also give a nearly-matching Ω(S2AH2/ϵ2)\Omega(S^2AH^2/\epsilon^2) lower bound, demonstrating the near-optimality of our algorithm in this setting.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext d269d478-adfd-4224-8a7d-411f1099b8d9

引用它的顶会 Paper139

问问它们各自怎么用它

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖