Detecting Data Deviations in Electronic Health Records
Kaiping Zheng, Horng Ruey Chua, Beng Chin Ooi
Abstract
Data deviations in electronic health records (EHR) refer to discrepancies between recorded entries and a patient's actual physiological state, indicating a decline in EHR data fidelity. Such deviations can result from pre-analytical variability, documentation errors, or unvalidated data sources. Effectively detecting data deviations is clinically valuable for identifying erroneous records, excluding them from downstream clinical workflows, and informing corrective actions. Despite its importance and practical relevance, this problem remains largely underexplored in existing research. To bridge this gap, we propose a bi-level knowledge distillation approach centered on a task-agnostic formulation of EHR data fidelity as an intrinsic measure of data reliability. Our approach performs layered knowledge distillation in two levels: from a computation-intensive, task-specific data Shapley oracle to a neural oracle for individual tasks, and then to a unified EHR data fidelity predictor. This design enables the integration of task-specific insights into a holistic assessment of a patient's EHR data fidelity from a multi-task perspective. By tracking the outputs of this learned predictor, we detect potential data deviations in EHR data. Experiments on both real-world EHR data from National University Hospital in Singapore and the public MIMIC-III dataset consistently validate the effectiveness of our approach in detecting data deviations in EHR data. Case studies further demonstrate its practical value in identifying clinically meaningful data deviations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8b93a97-4c63-467b-8698-49aaa8327882Builds on25
- FastSHAP: Real-Time Shapley Value EstimationNeil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee et al.ICLR 2022 · 186 citations
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 152 citations
- StageNet: Stage-Aware Neural Networks for Health Risk PredictionJunyi Gao, Cao Xiao, Yasha Wang, Wen Tang et al.WWW 2020 · 131 citations
- Validation Free and Replication Robust Volume-based Data ValuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang LowNeurIPS 2021 · 89 citations
- GRASP: Generic Framework for Health Status Representation Learning Based on Incorporating Knowledge from Similar PatientsChaohe Zhang, Xin Gao, Liantao Ma, Yasha Wang et al.AAAI 2021 · 77 citations
Related papers
- CLEAR: Addressing Representation Contamination in Multimodal Healthcare AnalyticsGe Su, Kaiping Zheng, Tiancheng Zhao, Jianwei YinKDD 2025 · 1 citation
- DeepAlerts: Deep Learning Based Multi-Horizon Alerts for Clinical Deterioration on Oncology Hospital WardsDingwen Li, Patrick G. Lyons, Chenyang Lu, Marin KollefAAAI 2020 · 18 citations
- FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare PredictionMuhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng et al.KDD 2024 · 5 citations
- SeqCare: Sequential Training with External Medical Knowledge Graph for Diagnosis Prediction in Healthcare DataYongxin Xu, Xu Chu, Kai Yang, Zhiyuan Wang et al.WWW 2023 · 41 citations
- DrFuse: Learning Disentangled Representation for Clinical Multi-Modal Fusion with Missing Modality and Modal InconsistencyWenfang Yao, Kejing Yin, William K. Cheung, Jia Liu et al.AAAI 2024 · 80 citations
