Counterfactual harm
Jonathan G. Richens, Rory Beard, Daniel H. Thompson
摘要
To act safely and ethically in the real world, agents must be able to reason about harm and avoid harmful actions. However, to date there is no statistical method for measuring harm and factoring it into algorithmic decisions. In this paper we propose the first formal definition of harm and benefit using causal models. We show that any factual definition of harm is incapable of identifying harmful actions in certain scenarios, and show that standard machine learning algorithms that cannot perform counterfactual reasoning are guaranteed to pursue harmful policies following distributional shifts. We use our definition of harm to devise a framework for harm-averse decision making using counterfactual objective functions. We demonstrate this framework on the problem of identifying optimal drug doses using a dose-response model learned from randomized control trial data. We find that the standard method of selecting doses using treatment effects results in unnecessarily harmful doses, while our counterfactual approach identifies doses that are significantly less harmful without sacrificing efficacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Counterfactual Identifiability of Bijective Causal ModelsArash Nasr-Esfahany, Mohammad Alizadeh, Devavrat ShahICML 2023 · 被引用 42 次
- Trustworthy Policy Learning under the Counterfactual No-Harm CriterionHaoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng 等ICML 2023 · 被引用 34 次
- Finding Counterfactually Optimal Action Sequences in Continuous State SpacesStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2023 · 被引用 18 次
- Partial Counterfactual Identification of Continuous Outcomes with a Curvature Sensitivity ModelValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelNeurIPS 2023 · 被引用 15 次
- Controlling Counterfactual Harm in Decision Support Systems Based on Prediction SetsEleni Straitouri, Suhas Thejaswi, Manuel Gomez RodriguezNeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper10
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Deep Structural Causal Models for Tractable Counterfactual InferenceNick Pawlowski, Daniel Coelho de Castro, Ben GlockerNeurIPS 2020 · 被引用 353 次
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 被引用 120 次
- Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for TextNishtha Madaan, Inkit Padhi, Naveen Panwar, Diptikalyan SahaAAAI 2021 · 被引用 115 次
- A Calculus for Stochastic Interventions: Causal Effect Identification and Surrogate ExperimentsJuan D. Correa, Elias BareinboimAAAI 2020 · 被引用 90 次
相关 Paper
- Learning Counterfactual Representations for Estimating Individual Dose-Response CurvesPatrick Schwab, Lorenz Linhardt, Stefan Bauer, Joachim M. Buhmann 等AAAI 2020 · 被引用 159 次
- Moral Responsibility for AI SystemsSander BeckersNeurIPS 2023 · 被引用 7 次
- A Causal Analysis of HarmSander Beckers, Hana Chockler, Joseph Y. HalpernNeurIPS 2022 · 被引用 25 次
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 被引用 60 次
- Generalization Bounds for Estimating Causal Effects of Continuous TreatmentsXin Wang, Shengfei Lyu, Xingyu Wu, Tianhao Wu 等NeurIPS 2022 · 被引用 38 次
