A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models
Jiayi Wang, Zhengling Qi, Raymond K. W. Wong
摘要
In this paper, we delve into the statistical analysis of the fitted Q-evaluation (FQE) method, which focuses on estimating the value of a target policy using offline data generated by some behavior policy. We provide a comprehensive theoretical understanding of FQE estimators under both parameteric and nonparametric models on the -function. Specifically, we address three key questions related to FQE that remain largely unexplored in the current literature: (1) Is the optimal convergence rate for estimating the policy value regarding the sample size () achievable for FQE under a non-parametric model with a fixed horizon ()? (2) How does the error bound depend on the horizon ? (3) What is the role of the probability ratio function in improving the convergence of FQE estimators? Specifically, we show that under the completeness assumption of -functions, which is mild in the non-parametric setting, the estimation errors for policy value using both parametric and non-parametric FQE estimators can achieve an optimal rate in terms of . The corresponding error bounds in terms of both and are also established. With an additional realizability assumption on ratio functions, the rate of estimation errors can be improved from to , which matches the sharpest known bound in the current literature under the tabular setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPsQi Kuang, Jiayi Wang, Fan Zhou, Zhengling QiNeurIPS 2025 · 被引用 3 次
- Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at RandomZiheng Wei, Annie Qu, Rui MiaoICML 2026
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
它引用的顶会 Paper5
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 被引用 161 次
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker 等ICLR 2021 · 被引用 112 次
- Variance-Aware Off-Policy Evaluation with Linear Function ApproximationYifei Min, Tianhao Wang, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 被引用 43 次
- On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy EvaluationXiaohong Chen, Zhengling QiICML 2022 · 被引用 36 次
- Sample Complexity of Nonparametric Off-Policy Evaluation on Low-Dimensional Manifolds using Deep NetworksXiang Ji, Minshuo Chen, Mengdi Wang, Tuo ZhaoICLR 2023 · 被引用 1 次
相关 Paper
- Bootstrapping Fitted Q-Evaluation for Off-Policy InferenceBotao Hao, Xiang Ji, Yaqi Duan, Hao Lu 等ICML 2021 · 被引用 46 次
- Off-Policy Fitted Q-Evaluation with Differentiable Function Approximators: Z-Estimation and Inference TheoryRuiqi Zhang, Xuezhou Zhang, Chengzhuo Ni, Mengdi WangICML 2022 · 被引用 20 次
- Learning Bellman Complete Representations for Offline Policy EvaluationJonathan D. Chang, Kaiwen Wang, Nathan Kallus, Wen SunICML 2022 · 被引用 18 次
- Optimal Estimation of Policy Gradient via Double Fitted IterationChengzhuo Ni, Ruiqi Zhang, Xiang Ji, Xuezhou Zhang 等ICML 2022 · 被引用 1 次
- On Gap-dependent Bounds for Offline Reinforcement LearningXinqi Wang, Qiwen Cui, Simon S. DuNeurIPS 2022 · 被引用 19 次
