When to Act and When to Ask: Policy Learning With Deferral Under Hidden Confounding
Marah Ghoummaid, Uri Shalit
Abstract
We consider the task of learning how to act in collaboration with a human expert based on observational data. The task is motivated by high-stake scenarios such as healthcare and welfare, where algorithmic action recommendations are made to a human expert, opening the option of deferring recommendation in cases where the human might act better on their own. This task is especially challenging when dealing with observational data, as using such data runs the risk of hidden confounders whose existence can lead to biased and harmful policies. However, unlike standard policy learning, the presence of a human expert can mitigate some of these risks. We build on the work of Mozannar and Sontag [2020] on consistent surrogate loss for learning with the option of deferral to an expert, where they solve a cost-sensitive supervised classification problem. Since we are solving a causal problem, where labels do not exist, we use a causal model to learn costs which are robust to a bounded degree of hidden confounding. We prove that our approach can take advantage of the strengths of both the model and the expert to obtain a better policy than either. We demonstrate our results by conducting experiments on synthetic and semi-synthetic data and show the advantages of our method compared to baselines.
Many works focus on solving the problem of policy learning for action recommendations from observational data. Some notable approaches include reweighting by inverse propensity weighting (IPW) and other weighting techniques [Swaminathan and Joachims, 2015, Kallus, 2017, Beygelzimer and Langford, 2009], and the approach of using doubly robust scores to determine the optimal treatment assignment policy for binary treatments [Dudík et al., 2014, Athey and Wager, 2021,?, Kallus and Zhou, 2020]. Other methods predict the Conditional Average Treatment Effect (CATE) and use it as the guideline for treatment assignment for each sample, such as Jesson et al. [2021], and Kallus et al. [2019]. Most of these works assume ignorability, i.e., that there are no hidden confounders that affect both treatment assignment and the outcome in the data. As mentioned earlier, this assumption rarely holds when observational data is in use, and its presence, if not accounted for, can lead to biased and harmful policies. Our work builds on previous work for learning supervised classification problems with the deferral option by Mozannar and Sontag [2020], and is inspired by Athey and Wager [2021] who derive costs for learning policies from observational data under the assumption of no hidden confounders. Gao and Yin [2023] present a framework for collaborative human-AI policy learning from observational data with deferral, building on earlier work [Gao et al., 2021] which did not allows for hidden confounding. To the best of our knowledge, theirs is the only existing method that learns a policy with the deferral option under hidden confounding. Their method minimizes an inverse propensity weighted estimator of the worst-case risk over a class of differentiable policies, and over an uncertainty set around the observed propensities. The uncertainty set is determined by constraints motivated by the Marginal Sensitivity Model [Tan, 2006]. While our method employs both outcome models and propensity scores, Gao and Yin [2023]'s approach focuses on propensity score re-weighting. The re-weighted objective implies that only cases where the proposed policy agrees to a high degree with the observed policy are taken into account. As we show in the experimental section
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 958fe6ed-2d9c-4557-bfbf-1757779260c0Cited by top-tier papers2
- A Causal Target for Learning to Defer Under Hidden ConfoundingYanmin Li, Lihua Liu, Xin Wang, Zhilong Mao et al.AAAI 2026
- Treatment Responder Classification with AbstentionHaoxiang Wang, Haoxuan Li, Ziyan Wang, Zhiheng Zhang et al.ICML 2026
Builds on5
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Quantifying Ignorance in Individual-Level Causal-Effect Estimates under Hidden ConfoundingAndrew Jesson, Sören Mindermann, Yarin Gal, Uri ShalitICML 2021 · 66 citations
- Sample Efficient Learning of Predictors that Complement HumansMohammad-Amin Charusaie, Hussein Mozannar, David A. Sontag, Samira SamadiICML 2022 · 52 citations
- B-Learner: Quasi-Oracle Bounds on Heterogeneous Causal Effects Under Hidden ConfoundingMiruna Oprescu, Jacob Dorn, Marah Ghoummaid, Andrew Jesson et al.ICML 2023 · 39 citations
- Fair Classifiers that Abstain without HarmTongxin Yin, Jean-Francois Ton, Ruocheng Guo, Yuanshun Yao et al.ICLR 2024 · 9 citations
Related papers
- Confounding-Robust Deferral Policy LearningRuijiang Gao, Mingzhang YinAAAI 2025 · 2 citations
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor et al.ICLR 2022 · 16 citations
- Exploiting Human-AI Dependence for Learning to DeferZixi Wei, Yuzhou Cao, Lei FengICML 2024 · 15 citations
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
- Online Decision MediationDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarNeurIPS 2022 · 5 citations
