GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
Xiang Lan, Feng Wu, Kai He, Qinghao Zhao, Shenda Hong, Mengling Feng
Abstract
While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between time series signals and visual ECG representations, and (2) limited explainability in linking diagnoses to granular waveform evidence. We introduce GEM, the first MLLM unifying ECG time series, 12-lead ECG images and text for grounded and clinician-aligned ECG interpretation. GEM enables feature-grounded analysis, evidence-driven reasoning, and a clinician-like diagnostic process through three core innovations: a dual-encoder framework extracting complementary time series and image features, cross-modal alignment for effective multimodal understanding, and knowledge-guided instruction generation for generating high-granularity grounding data (ECG-Grounding) linking diagnoses to measurable parameters (, QRS/PR Intervals). Additionally, we propose the Grounded ECG Understanding task, a clinically motivated benchmark designed to comprehensively assess the MLLM's capability in grounded ECG understanding. Experimental results on both existing and our proposed benchmarks show GEM significantly improves predictive performance (CSN ), explainability (), and grounding (), making it more suitable for real-world clinical applications. GitHub repository: https://github.com/lanxiang1017/GEM.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fc295d6-c02b-4522-b49e-e16af744575aCited by top-tier papers9
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO TrainingDavid Dai, Peilin Chen, Chanakya Ekbote, Paul Pu LiangNeurIPS 2025 · 48 citations
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language ModelsTong Guan, Zijie Meng, Dianqi Li, Shiyu Wang et al.ICLR 2026 · 29 citations
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-Wise Group Relative Policy OptimizationJingyi Zhang, Jiaxing Huang, Huanjin Yao, Shunyu Liu et al.ICCV 2025 · 17 citations
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and DiagnosisYuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue et al.CVPR 2026 · 11 citations
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang et al.NeurIPS 2025 · 3 citations
Builds on6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BIOT: Biosignal Transformer for Cross-data Learning in the WildChaoqi Yang, M. Brandon Westover, Jimeng SunNeurIPS 2023 · 345 citations
- Intra-Inter Subject Self-Supervised Learning for Multivariate Cardiac SignalsXiang Lan, Dianwen Ng, Shenda Hong, Mengling FengAAAI 2022 · 71 citations
- CLOCS: Contrastive Learning of Cardiac Signals Across Space, Time, and PatientsDani Kiyasseh, Tingting Zhu, David A. CliftonICML 2021 · 30 citations
- Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining ApproachXiang Lan, Hanshu Yan, Shenda Hong, Mengling FengICLR 2024 · 21 citations
Related papers
- ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG InterpretationJiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang et al.ICML 2026 · 11 citations
- HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic ReasoningJinning Yang, Wenjie Sun, Wen ShiAAAI 2026
- anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task UnderstandingHaitao Li, Ziyu Li, Yiheng Mao, Ziyi Liu et al.AAAI 2026 · 4 citations
- Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge EnhancementChe Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah et al.ICML 2024 · 83 citations
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisBo Liu, Ke Zou, Li-Ming Zhan, Zexin Lu et al.ICCV 2025 · 10 citations
