Lune

ICML2020顶会

Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound

Lin Yang, Mengdi Wang

2020年份
308被引次数
94顶会引用

摘要

Exploration in reinforcement learning (RL) suffers from the curse of dimensionality when the state-action space is large. A common practice is to parameterize the high-dimensional value and policy functions using given features. However existing methods either have no theoretical guarantee or suffer a regret that is exponential in the planning horizon HH. In this paper, we propose an online RL algorithm, namely the MatrixRL, that leverages ideas from linear bandit to learn a low-dimensional representation of the probability transition model while carefully balancing the exploitation-exploration tradeoff. We show that MatrixRL achieves a regret bound O(H2dlog⁡TT){O}\big(H^2d\log T\sqrt{T}\big) where dd is the number of features. MatrixRL has an equivalent kernelized version, which is able to work with an arbitrary kernel Hilbert space without using explicit features. In this case, the kernelized MatrixRL satisfies a regret bound O(H2d~log⁡TT){O}\big(H^2\widetilde{d}\log T\sqrt{T}\big), where d~\widetilde{d} is the effective dimension of the kernel space. To our best knowledge, for RL using features or kernels, our results are the first regret bounds that are near-optimal in time TT and dimension dd (or d~\widetilde{d}) and polynomial in the planning horizon HH.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 00e89180-5089-493d-a179-d02869ad99f3

引用它的顶会 Paper94

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖