Lune

ICML2024顶会

Eluder-based Regret for Stochastic Contextual MDPs

Orin Levy, Asaf B. Cassel, Alon Cohen, Yishay Mansour

2024年份
10被引次数
7顶会引用

摘要

We present the E-UC3^3RL algorithm for regret minimization in Stochastic Contextual Markov Decision Processes (CMDPs). The algorithm operates under the minimal assumptions of realizable function class and access to offline least squares and log loss regression oracles. Our algorithm is efficient (assuming efficient offline regression oracles) and enjoys a regret guarantee of O~(H3T∣S∣∣A∣dE(P)log⁡(∣F∣∣P∣/δ))),\widetilde{O}(H^3 \sqrt{T |S| |A|d_{\mathrm{E}}(\mathcal{P}) \log (|\mathcal{F}| |\mathcal{P}|/ \delta) )}) , with TT being the number of episodes, SS the state space, AA the action space, HH the horizon, P\mathcal{P} and F\mathcal{F} are finite function classes used to approximate the context-dependent dynamics and rewards, respectively, and dE(P)d_{\mathrm{E}}(\mathcal{P}) is the Eluder dimension of P\mathcal{P} w.r.t the Hellinger distance. To the best of our knowledge, our algorithm is the first efficient and rate-optimal regret minimization algorithm for CMDPs that operates under the general offline function approximation setting. In addition, we extend the Eluder dimension to general bounded metrics which may be of separate interest.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper7

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖