Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
Yuheng Zhang, Nan Jiang
Abstract
We investigate off-policy evaluation (OPE), a central and fundamental problem in reinforcement learning (RL), in the challenging setting of Partially Observable Markov Decision Processes (POMDPs) with large observation spaces. Recent works of Uehara et al. (2023a); Zhang & Jiang (2024) developed a modelfree framework and identified important coverage assumptions (called belief and outcome coverage) that enable accurate OPE of memoryless policies with polynomial sample complexities, but handling more general target policies that depend on the entire observable history remained an open problem. In this work, we prove information-theoretic hardness for model-free OPE of history-dependent policies in several settings, characterized by additional assumptions imposed on the behavior policy (memoryless vs. history-dependent) and/or the state-revealing property of the POMDP (single-step vs. multi-step revealing). We further show that some hardness can be circumvented by a natural model-based algorithm-whose analysis has surprisingly eluded the literature despite the algorithm's simplicitydemonstrating provable separation between model-free and model-based OPE in POMDPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8cc4d53-fb77-4dfb-bdc7-43a1098bc122Cited by top-tier papers2
- Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPsQi Kuang, Jiayi Wang, Fan Zhou, Zhengling QiNeurIPS 2025 · 3 citations
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
Builds on23
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro et al.NeurIPS 2021 · 339 citations
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain ClassifiersBenjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine et al.ICLR 2021 · 120 citations
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 91 citations
Related papers
- On the Curses of Future and History in Future-dependent Value Functions for Off-policy EvaluationYuheng Zhang, Nan JiangNeurIPS 2024 · 11 citations
- Model-based Reinforcement Learning for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICML 2024 · 5 citations
- Lower Bounds for Learning in Revealing POMDPsFan Chen, Huan Wang, Caiming Xiong, Song Mei et al.ICML 2023 · 18 citations
- Learning in Observable POMDPs, without Computationally Intractable OraclesNoah Golowich, Ankur Moitra, Dhruv RohatgiNeurIPS 2022 · 39 citations
- Planning and Learning in Partially Observable Systems via Filter StabilityNoah Golowich, Ankur Moitra, Dhruv RohatgiSTOC 2023 · 4 citations
