Trustworthy Policy Learning under the Counterfactual No-Harm Criterion
Haoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng, Yue Liu, Peng Wu
摘要
Trustworthy policy learning has significant importance in making reliable and harmless treatment decisions for individuals. Previous policy learning approaches aim at the well-being of subgroups by maximizing the utility function (e.g., conditional average causal effects, post-view click-through&conversion rate in recommendations), however, individual-level counterfactual no-harm criterion has rarely been discussed. In this paper, we first formalize the counterfactual no-harm criterion for policy learning from a principal stratification perspective. Next, we propose a novel upper bound for the fraction negatively affected by the policy and show the consistency and asymptotic normality of the estimator. Based on the estimators for the policy utility and harm upper bounds, we further propose a policy learning approach that satisfies the counterfactual noharm criterion, and prove its consistency to the optimal policy reward for parametric and nonparametric policy classes, respectively. Extensive experiments are conducted to show the effectiveness of the proposed policy learning approach for satisfying the counterfactual no-harm criterion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Optimal Transport for Treatment Effect EstimationHao Wang, Jiajun Fan, Zhichao Chen, Haoxuan Li 等NeurIPS 2023 · 被引用 71 次
- Propensity Matters: Measuring and Enhancing Balancing for RecommendationHaoxuan Li, Yanghao Xiao, Chunyuan Zheng, Peng Wu 等ICML 2023 · 被引用 55 次
- Policy Learning for Balancing Short-Term and Long-Term RewardsPeng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang 等ICML 2024 · 被引用 16 次
- Pareto Invariant Representation Learning for Multimedia RecommendationShanshan Huang, Haoxuan Li, Qingsong Li, Chunyuan Zheng 等ACM MM 2023 · 被引用 16 次
- Who Should Be Given Incentives? Counterfactual Optimal Treatment Regimes Learning for RecommendationHaoxuan Li, Chunyuan Zheng, Peng Wu, Kun Kuang 等KDD 2023 · 被引用 15 次
它引用的顶会 Paper4
- ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate EstimationHao Wang, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang 等SIGIR 2022 · 被引用 86 次
- What's the Harm? Sharp Bounds on the Fraction Negatively Affected by TreatmentNathan KallusNeurIPS 2022 · 被引用 40 次
- Unit Selection with Causal DiagramAng Li, Judea PearlAAAI 2022 · 被引用 25 次
- Counterfactual harmJonathan G. Richens, Rory Beard, Daniel H. ThompsonNeurIPS 2022 · 被引用 5 次
相关 Paper
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 被引用 10 次
- Practical Counterfactual Policy Learning for Top-K RecommendationsYaxu Liu, Jui-Nan Yen, Bo-Wen Yuan, Rundong Shi 等KDD 2022 · 被引用 11 次
- Fairness on Principal Stratum: A New Perspective on Counterfactual FairnessHaoxuan Li, Zeyu Tang, Zhichao Jiang, Zhuangyan Fang 等ICML 2025
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
- A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage OutcomesChenyang Li, Hao Mei, Yue LiuICML 2026
