Exploiting Negative Samples: A Catalyst for Cohort Discovery in Healthcare Analytics
Kaiping Zheng, Horng Ruey Chua, Melanie Herschel, H. V. Jagadish, Beng Chin Ooi, James Wei Luen Yip
摘要
In healthcare analytics, addressing binary diagnosis or prognosis tasks presents unique challenges due to the inherent asymmetry between positive and negative samples. While positive samples, indicating patients with a disease, are defined based on stringent medical criteria, negative samples are defined in an open-ended manner and remain underexplored in prior research. To bridge this gap, we propose an innovative approach to facilitate cohort discovery within negative samples, leveraging a Shapley-based exploration of interrelationships between these samples, which holds promise for uncovering valuable insights concerning the studied disease, and related comorbidity and complications. We quantify each sample's contribution using data Shapley values, subsequently constructing the Negative Sample Shapley Field to model the distribution of all negative samples. Next, we transform this field through manifold learning, preserving the essential data structure information while imposing an isotropy constraint in data Shapley values. Within this transformed space, we pinpoint cohorts of medical interest via density-based clustering. We empirically evaluate the effectiveness of our approach on the real-world electronic medical records from National University Hospital in Singapore, yielding clinically valuable insights aligned with existing knowledge, and benefiting medical research and clinical decision-making.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- CohortNet: Empowering Cohort Discovery for Interpretable Healthcare AnalyticsQingpeng Cai, Kaiping Zheng, H. V. Jagadish, Beng Chin Ooi 等VLDB 2024 · 被引用 10 次
- On the Impact of the Utility in Semivalue-based Data ValuationMélissa Tamine, Benjamin Heymann, Maxime Vono, Patrick LoiseauICLR 2026 · 被引用 3 次
- Detecting Data Deviations in Electronic Health RecordsKaiping Zheng, Horng Ruey Chua, Beng Chin OoiNeurIPS 2025 · 被引用 1 次
- NeuralCohort: Cohort-aware Neural Representation Learning for Healthcare AnalyticsChangshuo Liu, Lingze Zeng, Kaiping Zheng, Shaofeng Cai 等ICML 2025
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
它引用的顶会 Paper8
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
- Efficient nonparametric statistical inference on population feature importance using Shapley valuesBrian D. Williamson, Jean FengICML 2020 · 被引用 86 次
- Improving Medical Predictions by Irregular Multimodal Electronic Health Records ModelingXinlu Zhang, Shiyang Li, Zhiyu Chen, Xifeng Yan 等ICML 2023 · 被引用 54 次
- ARM-Net: Adaptive Relation Modeling Network for Structured DataShaofeng Cai, Kaiping Zheng, Gang Chen, H. V. Jagadish 等SIGMOD 2021 · 被引用 38 次
- UNITE: Uncertainty-based Health Risk Prediction Leveraging Multi-sourced DataChacha Chen, Junjie Liang, Fenglong Ma, Lucas Glass 等WWW 2021 · 被引用 29 次
相关 Paper
- A Comprehensive Study of Shapley Value in Data AnalyticsHong Lin, Shixin Wan, Zhongle Xie, Ke Chen 等VLDB 2025 · 被引用 4 次
- Tab-Shapley: Identifying Top-k Tabular Data Quality InsightsManisha Padala, Lokesh Nagalapatti, Atharv Tyagi, Ramasuri Narayanam 等AAAI 2025 · 被引用 1 次
- GRASP: Generic Framework for Health Status Representation Learning Based on Incorporating Knowledge from Similar PatientsChaohe Zhang, Xin Gao, Liantao Ma, Yasha Wang 等AAAI 2021 · 被引用 77 次
- Shapley explainability on the data manifoldChristopher Frye, Damien de Mijolla, Tom Begley, Laurence Cowton 等ICLR 2021 · 被引用 125 次
- Fast Shapley Value Computation in Data Assemblage Tasks as Cooperative Simple GamesXuan Luo, Jian Pei, Cheng Xu, Wenjie Zhang 等SIGMOD 2024 · 被引用 13 次
