Lune

ICML2021顶会

UCB Momentum Q-learning: Correcting the bias without forgetting

Pierre Ménard, Omar Darwiche Domingues, Xuedong Shang, Michal Valko

2021年份
53被引次数
36顶会引用

摘要

We propose UCBMQ, Upper Confidence Bound Momentum Q-learning, a new algorithm for reinforcement learning in tabular and possibly stagedependent, episodic Markov decision process. UCBMQ is based on Q-learning where we add a momentum term and rely on the principle of optimism in face of uncertainty to deal with exploration. Our new technical ingredient of UCBMQ is the use of momentum to correct the bias that Q-learning suffers while, at the same time, limiting the impact it has on the the second-order term of the regret. For UCBMQ, we are able to guarantee a regret of at most O( √ H 3 SAT + H 4 SA) where H is the length of an episode, S the number of states, A the number of actions, T the number of episodes and ignoring terms in poly log(SAHT ). Notably, UCBMQ is the first algorithm that simultaneously matches the lower bound of Ω( √ H 3 SAT ) for large enough T and has a second-order term (with respect to the horizon T ) that scales only linearly with the number of states S. 1 It is the same reason why there is an extra factor S in the first order term of the bound of UCRL algorithm by Jaksch et al. (2010) . This factor is "pushed" to the second-order term by the improved analysis of Azar et al. (2017) .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper36

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖
UCB Momentum Q-learning: Correcting the bias without forgetting | Lune Research