CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information Decoding
Zhihong Zhu, Yunyan Zhang, Fan Zhang, Bowen Xing, Xian Wu
Abstract
Medical Visual Question Answering (Med-VQA) aims to generate accurate answers for clinical questions grounded in medical images, which has attracted increasing research attention due to its potential to streamline diagnostics and reduce clinical burden. Recent advances in Large Vision-Language Models (LVLMs) have shown great promise for Med-VQA, but still suffer from two inference-time issues:
(1) attention shift, where the LVLM over-relies on textual priors; and (2) attention dispersion, where it fails to focus on critical diagnostic regions. To tackle these issues, we propose Contrastive Mutual Information Decoding (CMID), a training-free inference-time intervention grounded in information theory for Med-VQA. Concretely, CMID first identifies the Principal Focus Area (PFA) from decoder attention maps, then constructs focus-preserving and focus-excluding views to derive dual contrastive signals that simultaneously amplify salient visual cues and suppress background noise. Crucially, these corrective signals are adaptively scaled by a reliability-gated self-correction mechanism, based on the distributional shift induced by the PFA. Extensive experiments on three Med-VQA benchmarks demonstrate the effectiveness of CMID. Further analyses showcase its robust generalizability across diverse medical architectures and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b586ebd8-23e6-4a05-a91a-2ac9aceddbdeBuilds on13
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Medical Visual Question Answering via Conditional ReasoningLi-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen et al.ACM MM 2020 · 157 citations
- Caption-Aware Medical VQA via Semantic Focusing and Progressive Cross-Modality ComprehensionFu'ze Cong, Shibiao Xu, Li Guo, Yinbing TianACM MM 2022 · 34 citations
- MedCoT: Medical Chain of Thought via Hierarchical ExpertJiaxiang Liu, Yuan Wang, Jiawei Du, Joey Zhou et al.EMNLP 2024 · 15 citations
Related papers
- Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMsXiao Liang, Chenxi Liu, Zhi Ma, Di Wang et al.AAAI 2026 · 1 citation
- Cross-Modal Attention Calibration for LVLM Hallucination MitigationJiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma et al.CVPR 2026 · 23 citations
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsRuiying Peng, Xueyu Wu, Jing Lei, Lu Hou et al.CVPR 2026 · 4 citations
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive DecodingSicong Leng, Hang Zhang, Guanzheng Chen, Xin Li et al.CVPR 2024
- Enhancing Medical Large Vision-Language Models via Alignment DistillationAofei Chang, Ting Wang, Fenglong MaAAAI 2026
