Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
Yuanhao Zou, Zhaozheng Yin
摘要
Medical Visual Question Answering (Med-VQA) is a challenging task that requires a deep understanding of both medical images and textual questions. Although recent works leveraging Medical Vision-Language Pre-training (Med-VLP) have shown strong performance on the Med-VQA task, there is still no unified solution for modality alignment, and the issue of hard negatives remains underexplored. Additionally, commonly used knowledge fusion techniques for Med-VQA may introduce irrelevant information. In this work, we propose a framework to address these challenges through three key contributions: (1) a unified solution for heterogeneous modality alignments across multiple levels, modalities, views, and stages, leveraging methods like contrastive learning and optimal transport theory; (2) a hard negative mining method that employs soft labels for multi-modality alignments and enforces the hard negative pair discrimination; and (3) a Gated Cross-Attention Module for Med-VQA that integrates the answer vocabulary as prior knowledge and selects relevant information from it. Our framework outperforms the previous state-ofthe-art on widely used Med-VQA datasets like RAD-VQA, SLAKE, PathVQA and VQA-2019. The code is available at https://github.com/AlexCo1d/AMiF
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text UnderstandingJiayun Jin, Haolong Chai, Xueying Huang, Xiaoqing Guo 等CVPR 2026 · 被引用 4 次
- CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information DecodingZhihong Zhu, Yunyan Zhang, Fan Zhang, Bowen Xing 等AAAI 2026 · 被引用 1 次
- SNAPHARD CONTRAST LEARNINGChangpu Meng, Jie Yang, Wanqing Li, Yi GuoICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 被引用 624 次
相关 Paper
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionPeng Zhang, Zhihui Lai, Wenting Chen, Xu Wu 等AAAI 2026
- Align, Reason and Learn: Enhancing Medical Vision-and-Language Pre-training with KnowledgeZhihong Chen, Guanbin Li, Xiang WanACM MM 2022 · 被引用 82 次
- RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-trainingZheng Yuan, Qiao Jin, Chuanqi Tan, Zhengyun Zhao 等ACM MM 2023 · 被引用 33 次
- MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual RepresentationsZiyang Zhang, Yang Yu, Yucheng Chen, Xulei Yang 等CVPR 2025
- Enhancing Medical Large Vision-Language Models via Alignment DistillationAofei Chang, Ting Wang, Fenglong MaAAAI 2026
