On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy Evaluation
Xiaohong Chen, Zhengling Qi
摘要
We study the off-policy evaluation (OPE) problem in an infinite-horizon Markov decision process with continuous states and actions. We recast the Q -function estimation into a special form of the nonparametric instrumental variables (NPIV) estimation problem. We first show that under one mild condition the NPIV formulation of Q function estimation is well-posed in the sense of L 2 -measure of ill-posedness with respect to the data generating distribution, bypassing a strong assumption on the discount factor γ imposed in the recent literature for obtaining the L 2 convergence rates of various Q -function estimators. Thanks to this new well-posed property, we derive the first minimax lower bounds for the convergence rates of nonparametric estimation of Q -function and its derivatives in both sup-norm and L 2 -norm, which are shown to be the same as those for the classical nonparametric regression (Stone, 1982). We then propose a sieve two-stage least squares estimator and establish its rate-optimality in both norms under some mild conditions. Our general results on the well-posedness and the minimax lower bounds are of independent interest to study not only other nonparametric estimators for Q function but also efficient estimation on the value of any target policy in off-policy settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo 等ICML 2023 · 被引用 24 次
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou 等NeurIPS 2023 · 被引用 21 次
- Off-Policy Fitted Q-Evaluation with Differentiable Function Approximators: Z-Estimation and Inference TheoryRuiqi Zhang, Xuezhou Zhang, Chengzhuo Ni, Mengdi WangICML 2022 · 被引用 20 次
- When is Realizability Sufficient for Off-Policy Reinforcement Learning?Andrea ZanetteICML 2023 · 被引用 16 次
- Bellman Residual Orthogonalization for Offline Reinforcement LearningAndrea Zanette, Martin J. WainwrightNeurIPS 2022 · 被引用 14 次
它引用的顶会 Paper11
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro 等NeurIPS 2021 · 被引用 339 次
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 被引用 184 次
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 被引用 161 次
相关 Paper
- Off-Policy Evaluation for Episodic Partially Observable Markov Decision Processes under Non-Parametric ModelsRui Miao, Zhengling Qi, Xiaoke ZhangNeurIPS 2022 · 被引用 18 次
- Semiparametrically Efficient Off-Policy Evaluation in Linear Markov Decision ProcessesChuhan Xie, Wenhao Yang, Zhihua ZhangICML 2023 · 被引用 8 次
- Simultaneous Statistical Inference for Off-Policy Evaluation in Reinforcement LearningTianpai Luo, Xinyuan Fan, Weichi WuNeurIPS 2025 · 被引用 1 次
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov 等NeurIPS 2023 · 被引用 31 次
- A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision ProcessesChengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan JiangICML 2022 · 被引用 31 次
