Lune

NeurIPS2025顶会

Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPs

Qi Kuang, Jiayi Wang, Fan Zhou, Zhengling Qi

2025年份
3被引次数
1顶会引用

摘要

We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes (POMDPs) with unobserved confounding. Recent advances have introduced bridge-function to circumvent unmeasured confounding and develop estimators for the policy value, yet the statistical error bounds of them related to the length of horizon T and the size of the state-action space |O||A| remain largely unexplored. In this paper, we systematically investigate the finite-sample error bounds of OPE estimators in finite-horizon tabular confounded POMDPs. Specifically, we show that under certain rank conditions, the estimation error for policy value can achieve a rate of O ( T 1 . 5 / √ n ) , excluding the cardinality of the observation space |O| and the action space |A| . With an additional mild condition on the concentrability coefficients in confounded POMDPs, the rate of estimation error can be improved to O ( T/ √ n ) . We also show that for a fully history-dependent policy , the estimation error scales as O (cid:0) T/ √ n ( |O||A| ) T 2 (cid:1) , highlighting the exponential error dependence introduced by history-based proxies to infer hidden states. Furthermore, when the target policy is memoryless policy , the error bound improves to O (cid:0) T/ √ n (cid:112) |O||A| (cid:1) , which matches the optimal rate known for tabular MDPs. To the best of our knowledge, this is the first work to provide a comprehensive finite-sample analysis of OPE in confounded POMDPs.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖