Off-Policy Evaluation under Nonignorable Missing Data
Han Wang, Yang Xu, Wenbin Lu, Rui Song
Abstract
Off-Policy Evaluation (OPE) aims to estimate the value of a target policy using offline data collected from potentially different policies. In real-world applications, however, logged data often suffers from missingness. While OPE has been extensively studied in the literature, a theoretical understanding of how missing data affects OPE results remains unclear. In this paper, we investigate OPE in the presence of monotone missingness and theoretically demonstrate that the value estimates remain unbiased under ignorable missingness but can be biased under nonignorable (informative) missingness. To retain the consistency of value estimation, we propose an inverse probability weighting value estimator and conduct statistical inference to quantify the uncertainty of the estimates. Through a series of numerical experiments, we empirically demonstrate that our proposed estimator yields a more reliable value inference under missing data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb258504-dee3-4a62-b7b8-e0bdd145a402Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li et al.NeurIPS 2020 · 96 citations
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou et al.ICLR 2020 · 72 citations
- Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision MakingChengchun Shi, Runzhe Wan, Rui Song, Wenbin Lu et al.ICML 2020 · 45 citations
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 43 citations
Related papers
- Off-Policy Evaluation for Ranking Policies under Deterministic Logging PoliciesKoichi Tanaka, Kazuki Kawamura, Takanori Muroi, Yusuke Narita et al.ICLR 2026 · 1 citation
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
- Off-Policy Evaluation with Deficient Support Using Side InformationNicolò Felicioni, Maurizio Ferrari Dacrema, Marcello Restelli, Paolo CremonesiNeurIPS 2022 · 19 citations
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie et al.NeurIPS 2023 · 6 citations
- Partial Identification of Policy Values under Network InterferenceZiyan Wang, Yiran Liu, Zhiheng ZhangICML 2026 · 36 citations
