Lune

NeurIPS2024顶会

When to Act and When to Ask: Policy Learning With Deferral Under Hidden Confounding

Marah Ghoummaid, Uri Shalit

2024年份
4被引次数
2顶会引用

摘要

We consider the task of learning how to act in collaboration with a human expert based on observational data. The task is motivated by high-stake scenarios such as healthcare and welfare, where algorithmic action recommendations are made to a human expert, opening the option of deferring recommendation in cases where the human might act better on their own. This task is especially challenging when dealing with observational data, as using such data runs the risk of hidden confounders whose existence can lead to biased and harmful policies. However, unlike standard policy learning, the presence of a human expert can mitigate some of these risks. We build on the work of Mozannar and Sontag [2020] on consistent surrogate loss for learning with the option of deferral to an expert, where they solve a cost-sensitive supervised classification problem. Since we are solving a causal problem, where labels do not exist, we use a causal model to learn costs which are robust to a bounded degree of hidden confounding. We prove that our approach can take advantage of the strengths of both the model and the expert to obtain a better policy than either. We demonstrate our results by conducting experiments on synthetic and semi-synthetic data and show the advantages of our method compared to baselines.

Many works focus on solving the problem of policy learning for action recommendations from observational data. Some notable approaches include reweighting by inverse propensity weighting (IPW) and other weighting techniques [Swaminathan and Joachims, 2015, Kallus, 2017, Beygelzimer and Langford, 2009], and the approach of using doubly robust scores to determine the optimal treatment assignment policy for binary treatments [Dudík et al., 2014, Athey and Wager, 2021,?, Kallus and Zhou, 2020]. Other methods predict the Conditional Average Treatment Effect (CATE) and use it as the guideline for treatment assignment for each sample, such as Jesson et al. [2021], and Kallus et al. [2019]. Most of these works assume ignorability, i.e., that there are no hidden confounders that affect both treatment assignment and the outcome in the data. As mentioned earlier, this assumption rarely holds when observational data is in use, and its presence, if not accounted for, can lead to biased and harmful policies. Our work builds on previous work for learning supervised classification problems with the deferral option by Mozannar and Sontag [2020], and is inspired by Athey and Wager [2021] who derive costs for learning policies from observational data under the assumption of no hidden confounders. Gao and Yin [2023] present a framework for collaborative human-AI policy learning from observational data with deferral, building on earlier work [Gao et al., 2021] which did not allows for hidden confounding. To the best of our knowledge, theirs is the only existing method that learns a policy with the deferral option under hidden confounding. Their method minimizes an inverse propensity weighted estimator of the worst-case risk over a class of differentiable policies, and over an uncertainty set around the observed propensities. The uncertainty set is determined by constraints motivated by the Marginal Sensitivity Model [Tan, 2006]. While our method employs both outcome models and propensity scores, Gao and Yin [2023]'s approach focuses on propensity score re-weighting. The re-weighted objective implies that only cases where the proposed policy agrees to a high degree with the observed policy are taken into account. As we show in the experimental section

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖