Off-Policy Fitted Q-Evaluation with Differentiable Function Approximators: Z-Estimation and Inference Theory
Ruiqi Zhang, Xuezhou Zhang, Chengzhuo Ni, Mengdi Wang
摘要
Off-Policy Evaluation (OPE) serves as one of the cornerstones in Reinforcement Learning (RL). Fitted Q Evaluation (FQE) with various function approximators, especially deep neural networks, has gained practical success. While statistical analysis has proved FQE to be minimax-optimal with tabular, linear and several nonparametric function families, its practical performance with more general function approximator is less theoretically understood. We focus on FQE with general differentiable function approximators, making our theory applicable to neural function approximations. We approach this problem using the Z-estimation theory and establish the following results: The FQE estimation error is asymptotically normal with explicit variance determined jointly by the tangent space of the function class at the ground truth, the reward structure, and the distribution shift due to off-policy learning; The finite-sample FQE error bound is dominated by the same variance term, and it can also be bounded by function class-dependent divergence, which measures how the off-policy distribution shift intertwines with the function approximator. In addition, we study bootstrapping FQE estimators for error distribution inference and estimating confidence intervals, accompanied by a Cramer-Rao lower bound that matches our upper bounds. The Z-estimation analysis provides a generalizable theoretical framework for studying off-policy estimation in RL and provides sharp statistical theory for FQE with differentiable function approximators. Our Work FQE Yes General and Differentiable Completeness Asymptotic Normality, Cramer-Rao Lower Bound and Distributional Consistency Table 1: Comparison on Different Function Approximators of OPE
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Offline Reinforcement Learning with Closed-Form Policy Improvement OperatorsJiachen Li, Edwin Zhang, Ming Yin, Qinxun Bai 等ICML 2023 · 被引用 18 次
- Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline DataRuiqi Zhang, Andrea ZanetteNeurIPS 2023 · 被引用 12 次
- Non-stationary Reinforcement Learning under General Function ApproximationSongtao Feng, Ming Yin, Ruiquan Huang, Yu-Xiang Wang 等ICML 2023 · 被引用 11 次
- Sample Complexity of Nonparametric Off-Policy Evaluation on Low-Dimensional Manifolds using Deep NetworksXiang Ji, Minshuo Chen, Mengdi Wang, Tuo ZhaoICLR 2023 · 被引用 1 次
- Optimal Estimation of Policy Gradient via Double Fitted IterationChengzhuo Ni, Ruiqi Zhang, Xiang Ji, Xuezhou Zhang 等ICML 2022 · 被引用 1 次
它引用的顶会 Paper9
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 被引用 172 次
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 被引用 161 次
- Provably Efficient Reinforcement Learning for Discounted MDPs with Feature MappingDongruo Zhou, Jiafan He, Quanquan GuICML 2021 · 被引用 143 次
相关 Paper
- Bootstrapping Fitted Q-Evaluation for Off-Policy InferenceBotao Hao, Xiang Ji, Yaqi Duan, Hao Lu 等ICML 2021 · 被引用 46 次
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
- Distributional Offline Policy Evaluation with Predictive Error GuaranteesRunzhe Wu, Masatoshi Uehara, Wen SunICML 2023 · 被引用 19 次
- Offline Reinforcement Learning with Differentiable Function Approximation is Provably EfficientMing Yin, Mengdi Wang, Yu-Xiang WangICLR 2023
- A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric ModelsJiayi Wang, Zhengling Qi, Raymond K. W. WongICML 2024
