Lune

NeurIPS2022顶会

Computationally Efficient Horizon-Free Reinforcement Learning for Linear Mixture MDPs

Dongruo Zhou, Quanquan Gu

2022年份
60被引次数
23顶会引用

摘要

Recent studies have shown that episodic reinforcement learning (RL) is not more difficult than contextual bandits, even with a long planning horizon and unknown state transitions. However, these results are limited to either tabular Markov decision processes (MDPs) or computationally inefficient algorithms for linear mixture MDPs. In this paper, we propose the first computationally efficient horizon-free algorithm for linear mixture MDPs, which achieves the optimal O~(dK+d2)\tilde O(d\sqrt{K} +d^2) regret up to logarithmic factors. Our algorithm adapts a weighted least square estimator for the unknown transitional dynamic, where the weight is both variance-aware and uncertainty-aware. When applying our weighted least square estimator to heterogeneous linear bandits, we can obtain an O~(d∑k=1Kσk2+d)\tilde O(d\sqrt{\sum_{k=1}^K \sigma_k^2} +d) regret in the first KK rounds, where dd is the dimension of the context and σk2\sigma_k^2 is the variance of the reward in the kk-th round. This also improves upon the best-known algorithms in this setting when σk2\sigma_k^2's are known.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext d76b520e-df79-4fe8-908e-8735e423dfa2

引用它的顶会 Paper23

问问它们各自怎么用它

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖