From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining
Fuying Wang, Jiacheng Xu, Lequan Yu
Abstract
Electrocardiograms (ECGs) play a vital role in monitoring cardiac health and diagnosing heart diseases. However, traditional deep learning approaches for ECG analysis rely heavily on largescale manual annotations, which are both timeconsuming and resource-intensive to obtain. To overcome this limitation, self-supervised learning (SSL) has emerged as a promising alternative, enabling the extraction of robust ECG representations that can be efficiently transferred to various downstream tasks. While previous studies have explored SSL for ECG pretraining and multimodal ECG-language alignment, they often fail to capture the multi-scale nature of ECG signals. As a result, these methods struggle to learn generalized representations due to their inability to model the hierarchical structure of ECG data. To address this gap, we introduce MELP a novel Multi-scale ECG-Language Pretraining (MELP) model that fully leverages hierarchical supervision from ECG-text pairs. MELP first pretrains a cardiology-specific language model to enhance its understanding of clinical text. It then applies three levels of cross-modal supervision-at the token, beat, and rhythm levels-to align ECG signals with textual reports, capturing structured information across different time scales. We evaluate MELP on three public ECG datasets across multiple tasks, including zero-shot ECG classification, linear probing, and transfer learning. Experimental results demonstrate that MELP outperforms existing SSL methods, underscoring its effectiveness and adaptability across diverse clinical applications. Our code is available at https: //github.com/HKU-MedAI/MELP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Learning Cardiac Latent Representations in Vectorcardiogram SpaceBosong Huang, Panzhen Zhao, Zengxiang Li, Patricia Lee et al.ICML 2026
- Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation SpacesPratham Yashwante, Rose YuICML 2026
- SGERA: Stein-Guided ECG-Report Alignment for ECG Representation LearningJian Chen, Yipeng Du, Wenhao Yuan, Shuai Wang et al.ICML 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Reading Your Heart: Learning ECG Words and Sentences via Pre-training ECG Language ModelJiarui Jin, Haoyu Wang, Hongyan Li, Jun Li et al.ICLR 2025
- Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge EnhancementChe Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah et al.ICML 2024 · 83 citations
- Guiding Masked Representation Learning to Capture Spatio-Temporal Relationship of ElectrocardiogramYeongyeon Na, Minje Park, Yunwon Tae, Sunghoon JooICLR 2024 · 92 citations
- Tracing the Heart's Pathways: ECG Representation Learning from a Cardiac Conduction PerspectiveTan Pan, Yixuan Sun, Chen Jiang, Qiong Gao et al.AAAI 2026
- TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG DiagnosisXuewei Zhou, Yajie Meng, Pan Zeng, Xianfang Tang et al.CVPR 2026
