Domain constraints improve risk prediction when outcome data is missing
Sidhika Balachandar, Nikhil Garg, Emma Pierson
摘要
Machine learning models are often trained to predict the outcome resulting from a human decision. For example, if a doctor decides to test a patient for disease, will the patient test positive? A challenge is that historical decision-making determines whether the outcome is observed: we only observe test outcomes for patients doctors historically tested. Untested patients, for whom outcomes are unobserved, may differ from tested patients along observed and unobserved dimensions. We propose a Bayesian model class which captures this setting. The purpose of the model is to accurately estimate risk for both tested and untested patients. Estimating this model is challenging due to the wide range of possibilities for untested patients. To address this, we propose two domain constraints which are plausible in health settings: a prevalence constraint, where the overall disease prevalence is known, and an expertise constraint, where the human decision-maker deviates from purely risk-based decision-making only along a constrained feature set. We show theoretically and on synthetic data that domain constraints improve parameter inference. We apply our model to a case study of cancer risk prediction, showing that the model's inferred risk predicts cancer diagnoses, its inferred testing policy captures known public health policies, and it can identify suboptimalities in test allocation. Though our case study is in healthcare, our analysis reveals a general class of domain constraints which can improve model estimation in many settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Predictive Performance Comparison of Decision Policies Under ConfoundingLuke Guerdan, Amanda Coston, Ken Holstein, Steven WuICML 2024 · 被引用 1 次
- From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased DecisionsTrenton Chang, Jenna WiensICML 2024 · 被引用 1 次
它引用的顶会 Paper9
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- A Fine-Grained Analysis on Distribution ShiftOlivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi 等ICLR 2022 · 被引用 258 次
- Extending the WILDS Benchmark for Unsupervised AdaptationShiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao 等ICLR 2022 · 被引用 116 次
- Out-of-Domain Robustness via Targeted AugmentationsIrena Gao, Shiori Sagawa, Pang Wei Koh, Tatsunori Hashimoto 等ICML 2023 · 被引用 33 次
相关 Paper
- Defining Expertise: Applications to Treatment Effect EstimationAlihan Hüyük, Qiyao Wei, Alicia Curth, Mihaela van der SchaarICLR 2024 · 被引用 3 次
- Bayesian Inference for Correlated Human Experts and ClassifiersMarkelle Kelly, Alex James Boyd, Samuel Showalter, Mark Steyvers 等ICML 2025
- Statistical Inference Under Constrained Selection BiasSantiago Cortes-Gomez, Mateo Dulce-Rubio, Carlos Miguel Patiño, Bryan WilderICML 2024
- Incorporating Interpretable Output Constraints in Bayesian Neural NetworksWanqian Yang, Lars Lorch, Moritz A. Graule, Himabindu Lakkaraju 等NeurIPS 2020 · 被引用 17 次
- When Machine Learning Gets Personal: Evaluating Prediction and ExplanationLouisa Cornelis, Guillermo Bernardez, Haewon Jeong, Nina MiolaneICLR 2026
