Deep Contextual Clinical Prediction with Reverse Distillation
Rohan S. Kodialam, Rebecca Boiarsky, Justin Lim, Aditya Sai, Neil Dixit, David A. Sontag
Abstract
Healthcare providers are increasingly using machine learning to predict patient outcomes to make meaningful interventions. However, despite innovations in this area, deep learning models often struggle to match performance of shallow linear models in predicting these outcomes, making it difficult to leverage such techniques in practice. In this work, motivated by the task of clinical prediction from insurance claims, we present a new technique called reverse distillation which pretrains deep models by using high-performing linear models for initialization. We make use of the longitudinal structure of insurance claims datasets to develop Self Attention with Reverse Distillation, or SARD, an architecture that utilizes a combination of contextual embedding, temporal embedding and self-attention mechanisms and most critically is trained via reverse distillation. SARD outperforms state-of-the-art methods on multiple clinical prediction outcomes, with ablation studies revealing that reverse distillation is a primary driver of these improvements. Code is available at https://github.com/clinicalml/omop-learn .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on2
- ConCare: Personalized Clinical Feature Embedding via Capturing the Healthcare ContextLiantao Ma, Chaohe Zhang, Yasha Wang, Wenjie Ruan et al.AAAI 2020 · 190 citations
- StageNet: Stage-Aware Neural Networks for Health Risk PredictionJunyi Gao, Cao Xiao, Yasha Wang, Wen Tang et al.WWW 2020 · 131 citations
Related papers
- Fast, Accurate, and Simple Models for Tabular Data via Augmented DistillationRasool Fakoor, Jonas Mueller, Nick Erickson, Pratik Chaudhari et al.NeurIPS 2020 · 65 citations
- Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear ComplexityMutian He, Philip N. GarnerICLR 2025
- SeqCare: Sequential Training with External Medical Knowledge Graph for Diagnosis Prediction in Healthcare DataYongxin Xu, Xu Chu, Kai Yang, Zhiyuan Wang et al.WWW 2023 · 41 citations
- Distilling Knowledge from Publicly Available Online EMR Data to Emerging Epidemic for PrognosisLiantao Ma, Xinyu Ma, Junyi Gao, Xianfeng Jiao et al.WWW 2021 · 32 citations
- Reverse Distillation: Consistently Scaling Protein Language Model RepresentationsDarius Catrina, Christian Bepler, Samuel Sledzieski, Rohit SinghICLR 2026 · 2 citations
