Lune

ICML2024顶会

A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models

Jiayi Wang, Zhengling Qi, Raymond K. W. Wong

2024年份
3顶会引用

摘要

In this paper, we delve into the statistical analysis of the fitted Q-evaluation (FQE) method, which focuses on estimating the value of a target policy using offline data generated by some behavior policy. We provide a comprehensive theoretical understanding of FQE estimators under both parameteric and nonparametric models on the QQ-function. Specifically, we address three key questions related to FQE that remain largely unexplored in the current literature: (1) Is the optimal convergence rate for estimating the policy value regarding the sample size nn (n−1/2n^{-1/2}) achievable for FQE under a non-parametric model with a fixed horizon (TT)? (2) How does the error bound depend on the horizon TT? (3) What is the role of the probability ratio function in improving the convergence of FQE estimators? Specifically, we show that under the completeness assumption of QQ-functions, which is mild in the non-parametric setting, the estimation errors for policy value using both parametric and non-parametric FQE estimators can achieve an optimal rate in terms of nn. The corresponding error bounds in terms of both nn and TT are also established. With an additional realizability assumption on ratio functions, the rate of estimation errors can be improved from T1.5/nT^{1.5}/\sqrt{n} to T/nT/\sqrt{n}, which matches the sharpest known bound in the current literature under the tabular setting.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext d38b8a88-c33e-4644-bf74-ebcec32c0fee

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖