Off-Policy Evaluation under Nonignorable Missing Data
Han Wang, Yang Xu, Wenbin Lu, Rui Song
摘要
Off-Policy Evaluation (OPE) aims to estimate the value of a target policy using offline data collected from potentially different policies. In real-world applications, however, logged data often suffers from missingness. While OPE has been extensively studied in the literature, a theoretical understanding of how missing data affects OPE results remains unclear. In this paper, we investigate OPE in the presence of monotone missingness and theoretically demonstrate that the value estimates remain unbiased under ignorable missingness but can be biased under nonignorable (informative) missingness. To retain the consistency of value estimation, we propose an inverse probability weighting value estimator and conduct statistical inference to quantify the uncertainty of the estimates. Through a series of numerical experiments, we empirically demonstrate that our proposed estimator yields a more reliable value inference under missing data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li 等NeurIPS 2020 · 被引用 96 次
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou 等ICLR 2020 · 被引用 72 次
- Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision MakingChengchun Shi, Runzhe Wan, Rui Song, Wenbin Lu 等ICML 2020 · 被引用 45 次
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 被引用 43 次
相关 Paper
- Off-Policy Evaluation for Ranking Policies under Deterministic Logging PoliciesKoichi Tanaka, Kazuki Kawamura, Takanori Muroi, Yusuke Narita 等ICLR 2026 · 被引用 1 次
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito 等AAAI 2023 · 被引用 29 次
- Off-Policy Evaluation with Deficient Support Using Side InformationNicolò Felicioni, Maurizio Ferrari Dacrema, Marcello Restelli, Paolo CremonesiNeurIPS 2022 · 被引用 19 次
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie 等NeurIPS 2023 · 被引用 6 次
- Partial Identification of Policy Values under Network InterferenceZiyan Wang, Yiran Liu, Zhiheng ZhangICML 2026 · 被引用 36 次
