Lune

NeurIPS2022顶会

Reinforcement Learning in a Birth and Death Process: Breaking the Dependence on the State Space

Jonatha Anselmi, Bruno Gaujal, Louis-Sébastien Rebuffi

2022年份
3被引次数
1顶会引用

摘要

In this paper, we revisit the regret of undiscounted reinforcement learning in MDPs with a birth and death structure. Specifically, we consider a controlled queue with impatient jobs and the main objective is to optimize a trade-off between energy consumption and user-perceived performance. Within this setting, the diameter DD of the MDP is Ω(SS)\Omega(S^S), where SS is the number of states. Therefore, the existing lower and upper bounds on the regret at timeTT, of order O(DSAT)O(\sqrt{DSAT}) for MDPs with SS states and AA actions, may suggest that reinforcement learning is inefficient here. In our main result however, we exploit the structure of our MDPs to show that the regret of a slightly-tweaked version of the classical learning algorithm Ucrl2 is in fact upper bounded by O~(E2AT)\tilde{\mathcal{O}}(\sqrt{E_2AT}) where E2E_2 is related to the weighted second moment of the stationary measure of a reference policy. Importantly, E2E_2 is bounded independently of SS. Thus, our bound is asymptotically independent of the number of states and of the diameter. This result is based on a careful study of the number of visits performed by the learning algorithm to the states of the MDP, which is highly non-uniform.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖