Fine-grained Adaptive Visual Prompt for Generative Medical Visual Question Answering
Ting Yu, Zixuan Tong, Jun Yu, Ke Zhang
摘要
Medical Visual Question Answering (MedVQA) serves as an automated medical assistant, capable of answering patient queries and aiding physician diagnoses based on medical images and questions. Recent advancements have shown that incorporating Large Language Models (LLMs) into Med-VQA tasks significantly enhances the capability for answer generation. However, for tasks requiring fine-grained organlevel precise localization, relying solely on language prompts struggles to accurately locate relevant regions within medical images due to substantial background noise. To address this challenge, we explore the use of visual prompts in Med-VQA tasks for the first time and propose fine-grained adaptive visual prompts to enhance generative MedVQA. Specifically, we introduce an Adaptive Visual Prompt Creator that adaptively generates region-level visual prompts based on image characteristics of various organs, providing fine-grained references for LLMs during answer retrieval and generation from the medical domain, thereby improving the model's precise cross-modal localization capabilities on original images. Furthermore, we incorporate a Hierarchical Answer Generator with Parameter-Efficient Fine-Tuning (PEFT) techniques, significantly enhancing the model's understanding of spatial and contextual information with minimal parameter increase, promoting the alignment of representation learning with the medical space. Extensive experiments on VQA-RAD, SLAKE, and DME datasets validate the effectiveness of our proposed method, demonstrating its potential in generative MedVQA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic FrameworkYuexi Du, Jinglu Wang, Shujie Liu, Nicha C. Dvornek 等ICLR 2026 · 被引用 4 次
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 被引用 1 次
- Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question AnsweringMeihong Pan, Yefeng ZhengCVPR 2026
- Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain ModelingJiale Liu, Haoming Zhou, Yishu Liu, Bingzhi Chen 等AAAI 2026
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Visual Programming for Zero-Shot Open-Vocabulary 3D Visual GroundingZhihao Yuan, Jinke Ren, Chun-Mei Feng, Hengshuang Zhao 等CVPR 2024 · 被引用 19 次
相关 Paper
- LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundXuechen Guo, Wenhao Chai, Shiyan Li, Gaoang WangACM MM 2024 · 被引用 18 次
- Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language ModelsJiawei Chen, Dingkang Yang, Yue Jiang, Mingcheng Li 等ACM MM 2024 · 被引用 5 次
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 被引用 25 次
- Self-PT: Adaptive Self-Prompt Tuning for Low-Resource Visual Question AnsweringBowen Yuan, Sisi You, Bing-Kun BaoACM MM 2023 · 被引用 5 次
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisBo Liu, Ke Zou, Li-Ming Zhan, Zexin Lu 等ICCV 2025 · 被引用 10 次
