Lune

ICLR2023顶会

Learning Adversarial Linear Mixture Markov Decision Processes with Bandit Feedback and Unknown Transition

Canzhe Zhao, Ruofeng Yang, Baoxiang Wang, Shuai Li

出版方
2023年份
7顶会引用

摘要

We study reinforcement learning (RL) with linear function approximation, unknown transition, and adversarial losses in the bandit feedback setting. Specifically, the unknown transition probability function is a linear mixture model with a given feature mapping, and the learner only observes the losses of the experienced state-action pairs instead of the whole loss function. We propose an efficient algorithm LSUOB-REPS which achieves O~(dS2K+HSAK)\widetilde{O}(dS^2\sqrt{K}+\sqrt{HSAK}) regret guarantee with high probability, where dd is the ambient dimension of the feature mapping, SS is the size of the state space, AA is the size of the action space, HH is the episode length and KK is the number of episodes. Furthermore, we also prove a lower bound of order Ω(dHK+HSAK)\Omega(dH\sqrt{K}+\sqrt{HSAK}) for this setting. To the best of our knowledge, we make the first step to establish a provably efficient algorithm with a sublinear regret guarantee in this challenging setting and solve the open problem of .

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper7

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖