Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPs
Qi Kuang, Jiayi Wang, Fan Zhou, Zhengling Qi
Abstract
We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes (POMDPs) with unobserved confounding. Recent advances have introduced bridge-function to circumvent unmeasured confounding and develop estimators for the policy value, yet the statistical error bounds of them related to the length of horizon T and the size of the state-action space |O||A| remain largely unexplored. In this paper, we systematically investigate the finite-sample error bounds of OPE estimators in finite-horizon tabular confounded POMDPs. Specifically, we show that under certain rank conditions, the estimation error for policy value can achieve a rate of O ( T 1 . 5 / √ n ) , excluding the cardinality of the observation space |O| and the action space |A| . With an additional mild condition on the concentrability coefficients in confounded POMDPs, the rate of estimation error can be improved to O ( T/ √ n ) . We also show that for a fully history-dependent policy , the estimation error scales as O (cid:0) T/ √ n ( |O||A| ) T 2 (cid:1) , highlighting the exponential error dependence introduced by history-based proxies to infer hidden states. Furthermore, when the target policy is memoryless policy , the error bound improves to O (cid:0) T/ √ n (cid:112) |O||A| (cid:1) , which matches the optimal rate known for tabular MDPs. To the best of our knowledge, this is the first work to provide a comprehensive finite-sample analysis of OPE in confounded POMDPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4539efa1-8a8b-4940-8baf-b7d69a9c9c36Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 91 citations
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov et al.NeurIPS 2023 · 31 citations
- A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision ProcessesChengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan JiangICML 2022 · 31 citations
- Off-Policy Evaluation for Episodic Partially Observable Markov Decision Processes under Non-Parametric ModelsRui Miao, Zhengling Qi, Xiaoke ZhangNeurIPS 2022 · 18 citations
- On the Curses of Future and History in Future-dependent Value Functions for Off-policy EvaluationYuheng Zhang, Nan JiangNeurIPS 2024 · 11 citations
Related papers
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo et al.ICML 2023 · 24 citations
- Model-Free and Model-Based Policy Evaluation when Causality is UncertainDavid Bruns-SmithICML 2021 · 14 citations
- A Policy Gradient Method for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICLR 2024 · 5 citations
- Model-based Reinforcement Learning for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICML 2024 · 5 citations
- Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPsYuheng Zhang, Nan JiangICLR 2025
