Trustworthy Policy Learning under the Counterfactual No-Harm Criterion
Haoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng, Yue Liu, Peng Wu
Abstract
Trustworthy policy learning has significant importance in making reliable and harmless treatment decisions for individuals. Previous policy learning approaches aim at the well-being of subgroups by maximizing the utility function (e.g., conditional average causal effects, post-view click-through&conversion rate in recommendations), however, individual-level counterfactual no-harm criterion has rarely been discussed. In this paper, we first formalize the counterfactual no-harm criterion for policy learning from a principal stratification perspective. Next, we propose a novel upper bound for the fraction negatively affected by the policy and show the consistency and asymptotic normality of the estimator. Based on the estimators for the policy utility and harm upper bounds, we further propose a policy learning approach that satisfies the counterfactual noharm criterion, and prove its consistency to the optimal policy reward for parametric and nonparametric policy classes, respectively. Extensive experiments are conducted to show the effectiveness of the proposed policy learning approach for satisfying the counterfactual no-harm criterion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8a5f836-49ac-4795-b57f-4dc48e74c9f9Cited by top-tier papers17
- Optimal Transport for Treatment Effect EstimationHao Wang, Jiajun Fan, Zhichao Chen, Haoxuan Li et al.NeurIPS 2023 · 71 citations
- Propensity Matters: Measuring and Enhancing Balancing for RecommendationHaoxuan Li, Yanghao Xiao, Chunyuan Zheng, Peng Wu et al.ICML 2023 · 55 citations
- Policy Learning for Balancing Short-Term and Long-Term RewardsPeng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang et al.ICML 2024 · 16 citations
- Pareto Invariant Representation Learning for Multimedia RecommendationShanshan Huang, Haoxuan Li, Qingsong Li, Chunyuan Zheng et al.ACM MM 2023 · 16 citations
- Who Should Be Given Incentives? Counterfactual Optimal Treatment Regimes Learning for RecommendationHaoxuan Li, Chunyuan Zheng, Peng Wu, Kun Kuang et al.KDD 2023 · 15 citations
Builds on4
- ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate EstimationHao Wang, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang et al.SIGIR 2022 · 86 citations
- What's the Harm? Sharp Bounds on the Fraction Negatively Affected by TreatmentNathan KallusNeurIPS 2022 · 40 citations
- Unit Selection with Causal DiagramAng Li, Judea PearlAAAI 2022 · 25 citations
- Counterfactual harmJonathan G. Richens, Rory Beard, Daniel H. ThompsonNeurIPS 2022 · 5 citations
Related papers
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 10 citations
- Practical Counterfactual Policy Learning for Top-K RecommendationsYaxu Liu, Jui-Nan Yen, Bo-Wen Yuan, Rundong Shi et al.KDD 2022 · 11 citations
- Fairness on Principal Stratum: A New Perspective on Counterfactual FairnessHaoxuan Li, Zeyu Tang, Zhichao Jiang, Zhuangyan Fang et al.ICML 2025
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
- A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage OutcomesChenyang Li, Hao Mei, Yue LiuICML 2026
