GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
Xiang Lan, Feng Wu, Kai He, Qinghao Zhao, Shenda Hong, Mengling Feng
摘要
While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between time series signals and visual ECG representations, and (2) limited explainability in linking diagnoses to granular waveform evidence. We introduce GEM, the first MLLM unifying ECG time series, 12-lead ECG images and text for grounded and clinician-aligned ECG interpretation. GEM enables feature-grounded analysis, evidence-driven reasoning, and a clinician-like diagnostic process through three core innovations: a dual-encoder framework extracting complementary time series and image features, cross-modal alignment for effective multimodal understanding, and knowledge-guided instruction generation for generating high-granularity grounding data (ECG-Grounding) linking diagnoses to measurable parameters (, QRS/PR Intervals). Additionally, we propose the Grounded ECG Understanding task, a clinically motivated benchmark designed to comprehensively assess the MLLM's capability in grounded ECG understanding. Experimental results on both existing and our proposed benchmarks show GEM significantly improves predictive performance (CSN ), explainability (), and grounding (), making it more suitable for real-world clinical applications. GitHub repository: https://github.com/lanxiang1017/GEM.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO TrainingDavid Dai, Peilin Chen, Chanakya Ekbote, Paul Pu LiangNeurIPS 2025 · 被引用 48 次
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language ModelsTong Guan, Zijie Meng, Dianqi Li, Shiyu Wang 等ICLR 2026 · 被引用 29 次
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-Wise Group Relative Policy OptimizationJingyi Zhang, Jiaxing Huang, Huanjin Yao, Shunyu Liu 等ICCV 2025 · 被引用 17 次
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and DiagnosisYuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue 等CVPR 2026 · 被引用 11 次
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BIOT: Biosignal Transformer for Cross-data Learning in the WildChaoqi Yang, M. Brandon Westover, Jimeng SunNeurIPS 2023 · 被引用 345 次
- Intra-Inter Subject Self-Supervised Learning for Multivariate Cardiac SignalsXiang Lan, Dianwen Ng, Shenda Hong, Mengling FengAAAI 2022 · 被引用 71 次
- CLOCS: Contrastive Learning of Cardiac Signals Across Space, Time, and PatientsDani Kiyasseh, Tingting Zhu, David A. CliftonICML 2021 · 被引用 30 次
- Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining ApproachXiang Lan, Hanshu Yan, Shenda Hong, Mengling FengICLR 2024 · 被引用 21 次
相关 Paper
- ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG InterpretationJiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang 等ICML 2026 · 被引用 11 次
- HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic ReasoningJinning Yang, Wenjie Sun, Wen ShiAAAI 2026
- anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task UnderstandingHaitao Li, Ziyu Li, Yiheng Mao, Ziyi Liu 等AAAI 2026 · 被引用 4 次
- Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge EnhancementChe Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah 等ICML 2024 · 被引用 83 次
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisBo Liu, Ke Zou, Li-Ming Zhan, Zexin Lu 等ICCV 2025 · 被引用 10 次
