Contrastive Learning for Clinical Outcome Prediction with Partial Data Sources
Meng Xia, Jonathan Wilson, Benjamin Goldstein, Ricardo Henao
Abstract
The use of machine learning models to predict clinical outcomes from (longitudinal) electronic health record (EHR) data is becoming increasingly popular due to advances in deep architectures, representation learning, and the growing availability of large EHR datasets. Existing models generally assume access to the same data sources during both training and inference stages. However, this assumption is often challenged by the fact that real-world clinical datasets originate from various data sources (with distinct sets of covariates), which though can be available for training (in a research or retrospective setting), are more realistically only partially available (a subset of such sets) for inference when deployed. So motivated, we introduce Contrastive Learning for clinical Outcome Prediction with Partial data Sources (CLOPPS), that trains encoders to capture information across different data sources and then leverages them to build classifiers restricting access to a single data source. This approach can be used with existing cross-sectional or longitudinal outcome classification models. We present experiments on two real-world datasets demonstrating that CLOPPS consistently outperforms strong baselines in several practical scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- Exploring Accurate and Transparent Domain Adaptation in Predictive Healthcare via Concept-Grounded Orthogonal InferencePengfei Hu, Chang Lu, Feifan Liu, Yue NingICML 2026 · 1 citation
- Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality MissingnessZihan Liang, Ziwen Pan, Ruoxuan XiongEMNLP 2025
- CROCS: Clustering and Retrieval of Cardiac Signals Based on Patient Disease Class, Sex, and AgeDani Kiyasseh, Tingting Zhu, David A. CliftonNeurIPS 2021 · 10 citations
- LUMIN: A Longitudinal Multi-modal Knowledge Decomposition Network for Predicting Breast Cancer RecurrenceChunyao Lu, Tianyu Zhang, Xinglong Liang, Yuan Gao et al.AAAI 2026
- GRASP: Generic Framework for Health Status Representation Learning Based on Incorporating Knowledge from Similar PatientsChaohe Zhang, Xin Gao, Liantao Ma, Yasha Wang et al.AAAI 2021 · 77 citations
