Multiply Robust Off-policy Evaluation and Learning under Truncation by Death
Jianing Chu, Shu Yang, Wenbin Lu
摘要
Typical off-policy evaluation (OPE) and offpolicy learning (OPL) are not well-defined problems under "truncation by death", where the outcome of interest is not defined after some events, such as death. The standard OPE no longer yields consistent estimators, and the standard OPL results in suboptimal policies. In this paper, we formulate OPE and OPL using principal stratification under "truncation by death". We propose a survivor value function for a subpopulation whose outcomes are always defined regardless of treatment conditions. We establish a novel identification strategy under principal ignorability, and derive the semiparametric efficiency bound of an OPE estimator. Then, we propose multiply robust estimators for OPE and OPL. We show that the proposed estimators are consistent and asymptotically normal even with flexible semi/nonparametric models for nuisance functions approximation. Moreover, under mild rate conditions of nuisance functions approximation, the estimators achieve the semiparametric efficiency bound. Finally, we conduct experiments to demonstrate the empirical performance of the proposed estimators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Evaluating and Learning Optimal Dynamic Treatment Regimes under Truncation by DeathSihyung Park, Wenbin Lu, Shu YangNeurIPS 2025 · 被引用 1 次
- Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at RandomZiheng Wei, Annie Qu, Rui MiaoICML 2026
- Efficient Causal Decision Making with One-sided FeedbackJianing Chu, Shu Yang, Wenbin Lu, Pulak GhoshICLR 2025
- Off-Policy Evaluation under Nonignorable Missing DataHan Wang, Yang Xu, Wenbin Lu, Rui SongICML 2025
它引用的顶会 Paper1
相关 Paper
- Double Reinforcement Learning for Efficient and Robust Off-Policy EvaluationNathan Kallus, Masatoshi UeharaICML 2020 · 被引用 6 次
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
- Off-policy Evaluation Beyond Overlap: Sharp Partial Identification Under SmoothnessSamir Khan, Martin Saveski, Johan UganderICML 2024 · 被引用 4 次
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 被引用 44 次
- Semiparametrically Efficient Off-Policy Evaluation in Linear Markov Decision ProcessesChuhan Xie, Wenhao Yang, Zhihua ZhangICML 2023 · 被引用 8 次
