Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
Chong Ma, Hanqi Jiang, Wenting Chen, Yiwei Li, Zihao Wu, Xiaowei Yu, Zhengliang Liu, Lei Guo, Dajiang Zhu, Tuo Zhang, Dinggang Shen, Tianming Liu, Xiang Li
摘要
In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingTrong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti 等ICCV 2025
- A Generalized Label Shift Perspective for Cross-Domain Gaze EstimationHaoran Yang, Xiaohui Chen, Chuan-Xian RenNeurIPS 2025
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image RecognitionShih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena YeungICCV 2021 · 被引用 516 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
相关 Paper
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 被引用 1 次
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 被引用 40 次
- MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual InvariantChenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang 等CVPR 2024 · 被引用 20 次
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti 等NeurIPS 2022 · 被引用 302 次
- From Human Attention to Diagnosis: Semantic Patch-Level Integration of Vision-Language Models in Medical ImagingDmitry Lvov, Ilya PershinNeurIPS 2025 · 被引用 2 次
