Detecting Any instruction-to-answer interaction relationship: Universal Instruction-to-Answer Navigator for Med-VQA
Zhongze Wu, Hongyan Xu, Yitian Long, Shan You, Xiu Su, Jun Long, Yueyi Luo, Chang Xu
摘要
Abstract Medical Visual Question Answering (Med-VQA) interprets complex medical imagery using user instructions for precise diagnostics, yet faces challenges due to diverse, inadequately annotated images. In this paper, we introduce the Universal Instruction-Vision Navigator (Uni-Med) framework for extracting instruction-to-answer relationships, facilitating the understanding of visual evidence behind responses. Specifically, we design the Instruct-to-Answer Clues Interpreter (IAI) to generate visual explanations based on the answers and mark the core part of instructions with "real intent" labels. The IAI-Med VQA dataset, produced using IAI, is now publicly available to advance Med-VQA research. Additionally, our Token-Level Cut-Mix module dynamically aligns visual explanations with image patches, ensuring answers are traceable and learnable. We also implement intention-guided attention to minimize non-core instruction interference, sharpening focus on 'real intent'. Extensive experiments on SLAKE datasets show Uni-Med's superior accuracies ( 87.52 %closed, 86.12 % overall), outperforming MedVInT-PMC-VQA by 1.22 % and 0.92 %. Code and dataset are available at: https://github.com/zhongzee/Uni-Med-master .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware AlignmentGuoxin Zang, Xue Li, Donglin Di, Lanshun Nie 等ACM MM 2025 · 被引用 8 次
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningShibo Sun, Xue Li, Donglin Di, Mingjie Wei 等ACM MM 2025 · 被引用 4 次
- Intra-Image Mining and Symmetric Maximum Concept Matching for Few Shot Out-of-Distribution DetectionKaixiang Chen, Pengfei Fang, Hui XueAAAI 2026
- MARIS: Marine Open-Vocabulary Instance SegmentationBingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao 等CVPR 2026
- Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary DetectorsZhongze Wu, Xiu Su, Feng Yang, Dan Niu 等CVPR 2026
它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- Medical Visual Question Answering via Conditional ReasoningLi-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen 等ACM MM 2020 · 被引用 157 次
相关 Paper
- MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal OutputYanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan 等CVPR 2025
- Towards a Multimodal Large Language Model with Pixel-Level Insight for BiomedicineXiaoshuang Huang, Lingdong Shen, Jia Liu, Fangxin Shang 等AAAI 2025 · 被引用 32 次
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question AnsweringYuanhao Zou, Zhaozheng YinCVPR 2025
- MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsMohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi SoltanolkotabiICLR 2025
- MedCoT: Medical Chain of Thought via Hierarchical ExpertJiaxiang Liu, Yuan Wang, Jiawei Du, Joey Zhou 等EMNLP 2024 · 被引用 15 次
