A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage Outcomes
Chenyang Li, Hao Mei, Yue Liu
Abstract
When designing interventions to promote desired actions, two-stage agent heterogeneity -- encompassing both engagement with the intervention and completion of the desired action -- creates significant challenges in identifying optimal intervention policies. While this two-dimensional heterogeneity creates distinct agent response types with varying marginal policy returns, existing literature typically falls short in full identification of all agent types, leading to inefficient intervention allocations. To address the challenge of learning optimal policies that account for two-stage outcomes, we propose a minimax approach within a counterfactual principal strata framework. A value function, accommodating varying policy returns across six potentially non-identifiable principal strata, is designed and partially identified to minimize the worst-case value loss relative to three benchmark policies: never-treat, always-treat, and oracle. We introduce three estimators for optimal policy learning: Principal Outcome Regression (P-OR), Principal Inverse Propensity Scoring (P-IPS), and Principal Doubly Robust (P-DR), providing theoretical guarantees for their unbiasedness, robustness, and regret upper bounds. Extensive numerical experiments demonstrate the effectiveness and superiority of the proposed approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bd915f6-8052-48e1-ac87-2f5d69eca991Builds on9
- Trustworthy Policy Learning under the Counterfactual No-Harm CriterionHaoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng et al.ICML 2023 · 34 citations
- Algorithmic Decision Making with Conditional FairnessRenzhe Xu, Peng Cui, Kun Kuang, Bo Li et al.KDD 2020 · 26 citations
- Deep Global and Local Generative Model for RecommendationHuafeng Liu, Liping Jing, Jingxuan Wen, Zhicheng Wu et al.WWW 2020 · 24 citations
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Policy Learning for Balancing Short-Term and Long-Term RewardsPeng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang et al.ICML 2024 · 16 citations
Related papers
- Evaluating and Learning Optimal Dynamic Treatment Regimes under Truncation by DeathSihyung Park, Wenbin Lu, Shu YangNeurIPS 2025 · 1 citation
- Unified Minimax Optimization Framework for Propensity Score Estimation in Debiased RecommendationChunyuan Zheng, Haocheng Yang, Jinkun Chen, Shufeng Zhang et al.AAAI 2026 · 2 citations
- Multiply Robust Off-policy Evaluation and Learning under Truncation by DeathJianing Chu, Shu Yang, Wenbin LuICML 2023 · 6 citations
- Proximal Causal Learning of Conditional Average Treatment EffectsErik Sverdrup, Yifan CuiICML 2023 · 7 citations
- Treatment Responder Classification with AbstentionHaoxiang Wang, Haoxuan Li, Ziyan Wang, Zhiheng Zhang et al.ICML 2026
