Lune

ICML2022顶会

Nearly Optimal Policy Optimization with Stable at Any Time Guarantee

Tianhao Wu, Yunchang Yang, Han Zhong, Liwei Wang, Simon S. Du, Jiantao Jiao

2022年份
15被引次数
9顶会引用

摘要

Policy optimization methods are one of the most widely used classes of Reinforcement Learning (RL) algorithms. However, theoretical understanding of these methods remains insufficient. Even in the episodic (time-inhomogeneous) tabular setting, the state-of-the-art theoretical result of policy-based method in is only O~(S2AH4K)\tilde{O}(\sqrt{S^2AH^4K}) where SS is the number of states, AA is the number of actions, HH is the horizon, and KK is the number of episodes, and there is a SH\sqrt{SH} gap compared with the information theoretic lower bound Ω~(SAH3K)\tilde{\Omega}(\sqrt{SAH^3K}). To bridge such a gap, we propose a novel algorithm Reference-based Policy Optimization with Stable at Any Time guarantee (), which features the property"Stable at Any Time". We prove that our algorithm achieves O~(SAH3K+AH4K)\tilde{O}(\sqrt{SAH^3K} + \sqrt{AH^4K}) regret. When S>HS>H, our algorithm is minimax optimal when ignoring logarithmic factors. To our best knowledge, RPO-SAT is the first computationally efficient, nearly minimax optimal policy-based algorithm for tabular RL.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4c8c27b7-11bf-4f4c-b4ef-e544746eb34f

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖