Caption-Aware Medical VQA via Semantic Focusing and Progressive Cross-Modality Comprehension
Fu'ze Cong, Shibiao Xu, Li Guo, Yinbing Tian
Abstract
Medical Visual Question Answering as a specific-domain task requires substantive prior knowledge of medicine. However, deep learning techniques encounter severe problems of limited supervision due to the scarcity of well-annotated large-scale medical VQA datasets. As an alternative to facing the data limitation problem, image captioning can be introduced to learn summary information about the picture, which is beneficial to question answering. To this end, we propose a caption-aware VQA method that can read the summary information of image content and clinic diagnoses from plenty of medical images and answer the medical question with richer multimodality features. The proposed method consists of two novel components emphasizing semantic locations and semantic content respectively. Firstly, to extract and leverage the semantic locations implied in image captioning, similarity analysis is designed to summarize the attention maps generated from image captioning by their relevance and guide the visual model to focus on the semantic-rich regions. Besides, to combine the semantic content in the generated captions, we propose a Progressive Compact Bilinear Interactions structure to achieve cross-modality comprehension over the image, question and caption features by performing bilinear attention in a gradual manner. Qualitative and quantitative experiments on various medical datasets exhibit the superiority of the proposed approach compared to the state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Detecting Any instruction-to-answer interaction relationship: Universal Instruction-to-Answer Navigator for Med-VQAZhongze Wu, Hongyan Xu, Yitian Long, Shan You et al.ICML 2024 · 4 citations
- SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image SegmentationKe Yan, Qing Cai, Fan Zhang, Ziyan Cao et al.AAAI 2025 · 1 citation
- CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information DecodingZhihong Zhu, Yunyan Zhang, Fan Zhang, Bowen Xing et al.AAAI 2026 · 1 citation
- Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference OptimizationZixuan Huang, Zhihong Zhu, Xiaolong Liu, Yanchao Hao et al.ACL 2026
Related papers
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao et al.CVPR 2024
- Hierarchical Graph Attention Network for Few-shot Visual-Semantic LearningChengxiang Yin, Kun Wu, Zhengping Che, Bo Jiang et al.ICCV 2021 · 11 citations
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 35 citations
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
- MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQAHaowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren et al.CVPR 2026 · 2 citations
