Off-Policy Evaluation for Episodic Partially Observable Markov Decision Processes under Non-Parametric Models
Rui Miao, Zhengling Qi, Xiaoke Zhang
摘要
We study the problem of off-policy evaluation (OPE) for episodic Partially Observable Markov Decision Processes (POMDPs) with continuous states. Motivated by the recently proposed proximal causal inference framework, we develop a non-parametric identification result for estimating the policy value via a sequence of so-called V-bridge functions with the help of time-dependent proxy variables. We then develop a fitted-Q-evaluation-type algorithm to estimate V-bridge functions recursively, where a non-parametric instrumental variable (NPIV) problem is solved at each step. By analyzing this challenging sequential NPIV problem, we establish the finite-sample error bounds for estimating the V-bridge functions and accordingly that for evaluating the policy value, in terms of the sample size, length of horizon and so-called (local) measure of ill-posedness at each step. To the best of our knowledge, this is the first finite-sample error bound for OPE in POMDPs under non-parametric models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 被引用 19 次
- Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement LearningShuguang Yu, Shuxing Fang, Ruixin Peng, Zhengling Qi 等NeurIPS 2024 · 被引用 9 次
- Confounding Robust Deep Reinforcement Learning: A Causal ApproachMingxuan Li, Junzhe Zhang, Elias BareinboimNeurIPS 2025 · 被引用 7 次
- A Policy Gradient Method for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICLR 2024 · 被引用 5 次
- Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPsQi Kuang, Jiayi Wang, Fan Zhou, Zhengling QiNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper7
- Minimax Estimation of Conditional Moment ModelsNishanth Dikkala, Greg Lewis, Lester Mackey, Vasilis SyrgkanisNeurIPS 2020 · 被引用 125 次
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 被引用 91 次
- Sample-Efficient Reinforcement Learning of Undercomplete POMDPsChi Jin, Sham M. Kakade, Akshay Krishnamurthy, Qinghua LiuNeurIPS 2020 · 被引用 88 次
- Dual Instrumental Variable RegressionKrikamol Muandet, Arash Mehrjou, Si Kai Lee, Anant RajNeurIPS 2020 · 被引用 87 次
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 被引用 81 次
相关 Paper
- On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy EvaluationXiaohong Chen, Zhengling QiICML 2022 · 被引用 36 次
- A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision ProcessesChengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan JiangICML 2022 · 被引用 31 次
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov 等NeurIPS 2023 · 被引用 31 次
- Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at RandomZiheng Wei, Annie Qu, Rui MiaoICML 2026
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo 等ICML 2023 · 被引用 24 次
