TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG Diagnosis
Xuewei Zhou, Yajie Meng, Pan Zeng, Xianfang Tang, Feifei Cui, Qiangguo Jin, Jialiang Yang, Junlin Xu
Abstract
Cardiovascular disease (CVD) diagnosis relies heavily on electrocardiograms (ECGs). However, most existing selfsupervised uni-modal methods suffer from limited representational capacity, while multi-modal frameworks are hindered by coarse-grained semantic alignment across modalities, thus restricting their generalizability in clinical settings. To address these limitations, we propose TAMER, a Tri-modal contrastive Alignment and Multiscale Embedding Refinement framework that jointly models ECG recordings, spectrograms, and diagnostic reports. TAMER is composed of three key components: First, the trimodal feature encoding and projection (TFEP) module employs modality-specific encoders to extract global and local features from ECG recordings, spectrograms, and diagnostic reports, and projects them into latent spaces. Then, the global-local temporal-spectral alignment (GLTSA) module captures complementary rhythm-and wave-level characteristics via contrastive alignment and attentive interaction between temporal and spectral modalities. Finally, the reportaware alignment and refinement (RAAR) module performs diagnostic-level alignment and wave-level refinement with clinical reports, enabling semantic enrichment of ECG representations. Extensive experiments on three public ECG datasets demonstrate that TAMER achieves state-of-theart zero-shot classification performance (AUC: 81.2%) and strong cross-domain generalization (AUC: 83.1%), outperforming existing uni-modal and multi-modal baseline methods.The source code is available at https://github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c73776c8-916d-4997-b1bf-cf83283e4632Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge EnhancementChe Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah et al.ICML 2024 · 83 citations
- From Token to Rhythm: A Multi-Scale Approach for ECG-Language PretrainingFuying Wang, Jiacheng Xu, Lequan YuICML 2025
- Boosting Masked ECG-Text Auto-Encoders as Discriminative LearnersHung Manh Pham, Aaqib Saeed, Dong MaICML 2025
- Medical Vision-Language Pretraining with LLM-Guided Temporal SupervisionLiang Bai, Zhi Wang, Huimin Yan, Xian YangAAAI 2026
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu et al.ICCV 2023 · 47 citations
