Unlocking Cross-Modal Biosignal Synthesis: A Temporally-Aware VAE-Diffusion Model
Chenyang Xu, Dezhen Wang, Hao Wang
Abstract
Synthesizing authentic phonocardiograms (PCG) from ubiquitous electrocardiograms (ECG) is a critical task for accessible cardiac monitoring. Existing generative models, however, struggle to capture the heart's complex electromechanical coupling, failing to meet the dual requirements of temporal precision and physiological fidelity needed for clinically relevant waveform analysis. We introduce the Temporally-Aware VAE-Diffusion model, a synergistic hybrid architecture that resolves this trade-off. Our architecture enforces tight physiological coupling through an Enhanced Condition Fusion mechanism and explicitly models long-range cardiac dynamics via Temporal Attention Blocks. On the EPHNOGRAM benchmark, our model sets a new state-of-the-art, achieving a Pearson correlation of 0.9100.008, 95.95% S1 detection accuracy, and a precise 12.0 ms timing error, significantly outperforming leading diffusion and Transformer baselines. Crucially, our work provides a reproducible zero-shot transfer evaluation for ECG-to-PCG synthesis. Evaluated on the synchronized PhysioNet/CinC 2016 training-a/MITHSDB subset without target-domain training, our model preserves high waveform fidelity and clinically relevant timing structure under domain shift, including on pathological recordings. These results support cross-dataset robustness of the proposed synthesis framework, while downstream diagnostic validation remains an important direction for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b11c346-bd66-4bae-8376-47ecfb439928Builds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- SE-Diff: Simulator and Experience Enhanced Diffusion Model for Comprehensive ECG GenerationXiaoda Wang, Kaiqiao Han, Yuhao Xu, Xiao Luo et al.ICLR 2026 · 4 citations
- Region-Disentangled Diffusion Model for High-Fidelity PPG-to-ECG TranslationDebaditya Shome, Pritam Sarkar, Ali EtemadAAAI 2024 · 42 citations
- CardioGAN: Attentive Generative Adversarial Network with Dual Discriminators for Synthesis of ECG from PPGPritam Sarkar, Ali EtemadAAAI 2021 · 99 citations
- EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary FrameJiansong Zhang, Xiaying Yang, Xiaoling Luo, Linlin ShenCVPR 2026
- ME-GAN: Learning Panoptic Electrocardio Representations for Multi-view ECG Synthesis Conditioned on Heart DiseasesJintai Chen, Kuanlun Liao, Kun Wei, Haochao Ying et al.ICML 2022 · 29 citations
