When to Act and When to Ask: Policy Learning With Deferral Under Hidden Confounding
Marah Ghoummaid, Uri Shalit
摘要
We consider the task of learning how to act in collaboration with a human expert based on observational data. The task is motivated by high-stake scenarios such as healthcare and welfare, where algorithmic action recommendations are made to a human expert, opening the option of deferring recommendation in cases where the human might act better on their own. This task is especially challenging when dealing with observational data, as using such data runs the risk of hidden confounders whose existence can lead to biased and harmful policies. However, unlike standard policy learning, the presence of a human expert can mitigate some of these risks. We build on the work of Mozannar and Sontag [2020] on consistent surrogate loss for learning with the option of deferral to an expert, where they solve a cost-sensitive supervised classification problem. Since we are solving a causal problem, where labels do not exist, we use a causal model to learn costs which are robust to a bounded degree of hidden confounding. We prove that our approach can take advantage of the strengths of both the model and the expert to obtain a better policy than either. We demonstrate our results by conducting experiments on synthetic and semi-synthetic data and show the advantages of our method compared to baselines.
Many works focus on solving the problem of policy learning for action recommendations from observational data. Some notable approaches include reweighting by inverse propensity weighting (IPW) and other weighting techniques [Swaminathan and Joachims, 2015, Kallus, 2017, Beygelzimer and Langford, 2009], and the approach of using doubly robust scores to determine the optimal treatment assignment policy for binary treatments [Dudík et al., 2014, Athey and Wager, 2021,?, Kallus and Zhou, 2020]. Other methods predict the Conditional Average Treatment Effect (CATE) and use it as the guideline for treatment assignment for each sample, such as Jesson et al. [2021], and Kallus et al. [2019]. Most of these works assume ignorability, i.e., that there are no hidden confounders that affect both treatment assignment and the outcome in the data. As mentioned earlier, this assumption rarely holds when observational data is in use, and its presence, if not accounted for, can lead to biased and harmful policies. Our work builds on previous work for learning supervised classification problems with the deferral option by Mozannar and Sontag [2020], and is inspired by Athey and Wager [2021] who derive costs for learning policies from observational data under the assumption of no hidden confounders. Gao and Yin [2023] present a framework for collaborative human-AI policy learning from observational data with deferral, building on earlier work [Gao et al., 2021] which did not allows for hidden confounding. To the best of our knowledge, theirs is the only existing method that learns a policy with the deferral option under hidden confounding. Their method minimizes an inverse propensity weighted estimator of the worst-case risk over a class of differentiable policies, and over an uncertainty set around the observed propensities. The uncertainty set is determined by constraints motivated by the Marginal Sensitivity Model [Tan, 2006]. While our method employs both outcome models and propensity scores, Gao and Yin [2023]'s approach focuses on propensity score re-weighting. The re-weighted objective implies that only cases where the proposed policy agrees to a high degree with the observed policy are taken into account. As we show in the experimental section
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A Causal Target for Learning to Defer Under Hidden ConfoundingYanmin Li, Lihua Liu, Xin Wang, Zhilong Mao 等AAAI 2026
- Treatment Responder Classification with AbstentionHaoxiang Wang, Haoxuan Li, Ziyan Wang, Zhiheng Zhang 等ICML 2026
它引用的顶会 Paper5
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 被引用 267 次
- Quantifying Ignorance in Individual-Level Causal-Effect Estimates under Hidden ConfoundingAndrew Jesson, Sören Mindermann, Yarin Gal, Uri ShalitICML 2021 · 被引用 66 次
- Sample Efficient Learning of Predictors that Complement HumansMohammad-Amin Charusaie, Hussein Mozannar, David A. Sontag, Samira SamadiICML 2022 · 被引用 52 次
- B-Learner: Quasi-Oracle Bounds on Heterogeneous Causal Effects Under Hidden ConfoundingMiruna Oprescu, Jacob Dorn, Marah Ghoummaid, Andrew Jesson 等ICML 2023 · 被引用 39 次
- Fair Classifiers that Abstain without HarmTongxin Yin, Jean-Francois Ton, Ruocheng Guo, Yuanshun Yao 等ICLR 2024 · 被引用 9 次
相关 Paper
- Confounding-Robust Deferral Policy LearningRuijiang Gao, Mingzhang YinAAAI 2025 · 被引用 2 次
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor 等ICLR 2022 · 被引用 16 次
- Exploiting Human-AI Dependence for Learning to DeferZixi Wei, Yuzhou Cao, Lei FengICML 2024 · 被引用 15 次
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
- Online Decision MediationDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarNeurIPS 2022 · 被引用 5 次
