Lune

ICLR2023顶会

Achieving Sub-linear Regret in Infinite Horizon Average Reward Constrained MDP with Linear Function Approximation

Arnob Ghosh, Xingyu Zhou, Ness B. Shroff

出版方
2023年份
8顶会引用

摘要

We study the infinite horizon average reward constrained Markov Decision Process (CMDP). In contrast to existing works on model-based, finite state space, we consider the model-free linear CMDP setup. We first propose a computationally inefficient algorithm and show that O~(d3T)\tilde{\mathcal{O}}(\sqrt{d^3T}) regret and constraint violation can be achieved, in which TT is the number of interactions, and dd is the dimension of the feature mapping. We also propose an efficient variant based on the primal-dual adaptation of the LSVI-UCB algorithm and show that O~((dT)3/4)\tilde{\mathcal{O}}((dT)^{3/4}) regret and constraint violation can be achieved. This improves the known regret bound of O~(T5/6)\tilde{\mathcal{O}}(T^{5/6}) for the finite state-space model-free constrained RL which was obtained under a stronger assumption compared to ours. We also develop an efficient policy-based algorithm via novel adaptation of the MDP-EXP2 algorithm to our primal-dual set up with O~(T)\tilde{\mathcal{O}}(\sqrt{T}) regret and even zero constraint violation bound under a stronger set of assumptions.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖