Minimax-Optimal Off-Policy Evaluation with Linear Function Approximation
Yaqi Duan, Zeyu Jia, Mengdi Wang
Abstract
This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cumulative value of a new target policy from logged history generated by unknown behavioral policies. We study a regression-based fitted Q iteration method, and show that it is equivalent to a model-based method that estimates a conditional mean embedding of the transition operator. We prove that this method is information-theoretically optimal and has nearly minimal estimation error. In particular, by leveraging contraction property of Markov processes and martingale concentration, we establish a finite-sample instance-dependent error upper bound and a nearly-matching minimax lower bound. The policy evaluation error depends sharply on a restricted -divergence over the function class between the long-term distribution of the target policy and the distribution of past data. This restricted -divergence is both instance-dependent and function-class-dependent. It characterizes the statistical limit of off-policy evaluation. Further, we provide an easily computable confidence bound for the policy evaluator, which may be useful for optimistic planning and safe policy improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 009f90b9-18ec-494c-94a1-12e82f1ee799Cited by top-tier papers79
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro et al.NeurIPS 2021 · 339 citations
- Offline RL Without Off-Policy EvaluationDavid Brandfonbrener, Will Whitney, Rajesh Ranganath, Joan BrunaNeurIPS 2021 · 217 citations
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 176 citations
Related papers
- Sparse Feature Selection Makes Batch Reinforcement Learning More Sample EfficientBotao Hao, Yaqi Duan, Tor Lattimore, Csaba Szepesvári et al.ICML 2021 · 29 citations
- Bootstrapping Fitted Q-Evaluation for Off-Policy InferenceBotao Hao, Xiang Ji, Yaqi Duan, Hao Lu et al.ICML 2021 · 46 citations
- Off-Policy Fitted Q-Evaluation with Differentiable Function Approximators: Z-Estimation and Inference TheoryRuiqi Zhang, Xuezhou Zhang, Chengzhuo Ni, Mengdi WangICML 2022 · 20 citations
- Offline Reinforcement Learning with Differentiable Function Approximation is Provably EfficientMing Yin, Mengdi Wang, Yu-Xiang WangICLR 2023
- Variance-Aware Off-Policy Evaluation with Linear Function ApproximationYifei Min, Tianhao Wang, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 43 citations
