From Human Attention to Diagnosis: Semantic Patch-Level Integration of Vision-Language Models in Medical Imaging
Dmitry Lvov, Ilya Pershin
Abstract
Predicting human eye movements during goal-directed visual search is critical for enhancing interactive AI systems. In medical imaging, such prediction can support radiologists in interpreting complex data, such as chest X-rays. Many existing methods rely on generic vision–language models and saliency-based features, which can limit their ability to capture fine-grained clinical semantics and integrate domain knowledge effectively. We present LogitGaze-Med , a state-of-the-art multimodal transformer framework that unifies (1) domain-specific visual encoders (e.g., CheXNet), (2) textual embeddings of diagnostic labels, and (3) semantic priors extracted via the logit-lens from an instruction-tuned medical vision–language model (LLaVA-Med). By directly predicting continuous fixation coordinates and dwell durations, our model generates clinically meaningful scanpaths. Experiments on the GazeSearch dataset and synthetic scanpaths generated from MIMIC-CXR and validated by experts demonstrate that LogitGaze-Med improves scanpath similarity metrics by 20–30% over competitive baselines and yields over 5% gains in downstream pathology classification when incorporating predicted fixations as additional training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8268c48-eca4-4d3a-915b-455f8d9c42f3Builds on6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang et al.NeurIPS 2024 · 43 citations
- Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized DataYucheng Shi, Quanzheng Li, Jin Sun, Xiang Li et al.ICLR 2025
- Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human AttentionSounak Mondal, Zhibo Yang, Seoyoung Ahn, Dimitris Samaras et al.CVPR 2023
- Towards Interpreting Visual Information Processing in Vision-Language ModelsClement Neo, Luke Ong, Philip Torr, Mor Geva et al.ICLR 2025
Related papers
- Interpreting Radiologist's Intention from Eye Movements in Chest X-ray DiagnosisTrong-Thang Pham, Anh Nguyen, Zhigang Deng, Carol C. Wu et al.ACM MM 2025 · 1 citation
- Rethinking Radiology Report Generation: From Narrative Flow to Topic-Guided FindingsSheng Cheng, Devika SubramanianICLR 2026
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingTrong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti et al.ICCV 2025
- LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic PathologiesAmeer Hamza, Abdullah, Yong Hyun Ahn, Sungyoung Lee et al.AAAI 2025 · 14 citations
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa et al.CHI 2024 · 81 citations
