A Flexible Generative Model for Heterogeneous Tabular EHR with Missing Modality
Huan He, William Hao, Yuanzhe Xi, Yong Chen, Bradley A. Malin, Joyce C. Ho
摘要
Realistic synthetic electronic health records (EHRs) can be leveraged to accelerate methodological developments for research purposes while mitigating privacy concerns associated with data sharing. However, the training of Generative Adversarial Networks remains challenging, often resulting in issues like mode collapse. While diffusion models have demonstrated progress in generating quality synthetic samples for tabular EHRs given ample denoising steps, their performance wanes when confronted with missing modalities in heterogeneous tabular EHRs data. For example, some EHRs contain solely static measurements, and some contain only contain temporal measurements, or a blend of both data types. To bridge this gap, we introduce FLEXGEN-EHR-a versatile diffusion model tailored for heterogeneous tabular EHRs, equipped with the capability of handling missing modalities in an integrative learning framework. We define an optimal transport module to align and accentuate the common feature space of heterogeneity of EHRs. We empirically show that our model consistently outperforms existing state-of-the-art synthetic EHR generation methods both in fidelity by up to 3.10% and utility by up to 7.16%. Additionally, we show that our method can be successfully used in privacy-sensitive settings, where the original patient-level data cannot be shared.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality MissingnessZihan Liang, Ziwen Pan, Ruoxuan XiongEMNLP 2025
- Generating Multi-Table Time Series EHR from Latent Space with Minimal PreprocessingEunbyeol Cho, Jiyoun Kim, Minjae Lee, Sungjin Park 等NeurIPS 2025
它引用的顶会 Paper4
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 被引用 903 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- WaveGrad: Estimating Gradients for Waveform GenerationNanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss 等ICLR 2021 · 被引用 44 次
相关 Paper
- MedDiTPro: A Prompt-Guided Diffusion Transformer for Multimodal Longitudinal Medical Data SynthesisYuan Zhong, Xiaochen Wang, Jiaqi Wang, Xiaokun Zhang 等KDD 2025
- Synthesizing Multimodal Electronic Health Records via Predictive Diffusion ModelsYuan Zhong, Xiaochen Wang, Jiaqi Wang, Xiaokun Zhang 等KDD 2024 · 被引用 4 次
- IGAMT: Privacy-Preserving Electronic Health Record Synthesization with Heterogeneity and IrregularityWenjie Wang, Pengfei Tang, Jian Lou, Yuanming Shao 等AAAI 2024 · 被引用 7 次
- PromptEHR: Conditional Electronic Healthcare Records Generation with Prompt LearningZifeng Wang, Jimeng SunEMNLP 2022 · 被引用 19 次
- Addressing Asynchronicity in Clinical Multimodal Fusion via Individualized Chest X-ray GenerationWenfang Yao, Chen Liu, Kejing Yin, William Kwok-Wai Cheung 等NeurIPS 2024 · 被引用 11 次
