Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
Yuheng Zhang, Nan Jiang
摘要
We investigate off-policy evaluation (OPE), a central and fundamental problem in reinforcement learning (RL), in the challenging setting of Partially Observable Markov Decision Processes (POMDPs) with large observation spaces. Recent works of Uehara et al. (2023a); Zhang & Jiang (2024) developed a modelfree framework and identified important coverage assumptions (called belief and outcome coverage) that enable accurate OPE of memoryless policies with polynomial sample complexities, but handling more general target policies that depend on the entire observable history remained an open problem. In this work, we prove information-theoretic hardness for model-free OPE of history-dependent policies in several settings, characterized by additional assumptions imposed on the behavior policy (memoryless vs. history-dependent) and/or the state-revealing property of the POMDP (single-step vs. multi-step revealing). We further show that some hardness can be circumvented by a natural model-based algorithm-whose analysis has surprisingly eluded the literature despite the algorithm's simplicitydemonstrating provable separation between model-free and model-based OPE in POMDPs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Breaking the Order Barrier: Off-Policy Evaluation for Confounded POMDPsQi Kuang, Jiayi Wang, Fan Zhou, Zhengling QiNeurIPS 2025 · 被引用 3 次
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
它引用的顶会 Paper23
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro 等NeurIPS 2021 · 被引用 339 次
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain ClassifiersBenjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine 等ICLR 2021 · 被引用 120 次
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 被引用 91 次
相关 Paper
- On the Curses of Future and History in Future-dependent Value Functions for Off-policy EvaluationYuheng Zhang, Nan JiangNeurIPS 2024 · 被引用 11 次
- Model-based Reinforcement Learning for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICML 2024 · 被引用 5 次
- Lower Bounds for Learning in Revealing POMDPsFan Chen, Huan Wang, Caiming Xiong, Song Mei 等ICML 2023 · 被引用 18 次
- Learning in Observable POMDPs, without Computationally Intractable OraclesNoah Golowich, Ankur Moitra, Dhruv RohatgiNeurIPS 2022 · 被引用 39 次
- Planning and Learning in Partially Observable Systems via Filter StabilityNoah Golowich, Ankur Moitra, Dhruv RohatgiSTOC 2023 · 被引用 4 次
