Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
Aofei Chang, Le Huang, Alex James Boyd, Parminder Bhatia, Taha A. Kass-Hout, Cao Xiao, Fenglong Ma
摘要
Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A 3 TUNE, a novel fine-tuning framework for Automatic Attention Alignment Tuning. A 3 TUNE leverages zeroshot weak labels from SAM, refines them into prompt-aware labels using BiomedCLIP, and then selectively modifies visually-critical attention heads to improve alignment while minimizing interference. Additionally, we introduce a A 3 MOE module, enabling adaptive parameter selection for attention tuning across diverse prompts and images. Extensive experiments on medical VQA and report generation benchmarks show that A 3 TUNE outperforms state-ofthe-art baselines, achieving enhanced attention distributions and performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Enhancing Medical Large Vision-Language Models via Alignment DistillationAofei Chang, Ting Wang, Fenglong MaAAAI 2026
- Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference OptimizationZixuan Huang, Zhihong Zhu, Xiaolong Liu, Yanchao Hao 等ACL 2026
- MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language ModelsAofei Chang, Le Huang, Alex Boyd, parminder bhatia 等ICML 2026
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical ImagesGuimeng Liu, Tianze Yu, Somayeh Ebrahimkhani, Lin Zhi Zheng Shawn 等ICLR 2026 · 被引用 3 次
- ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language ModelsJunzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu 等CVPR 2025
- Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language ModelsJia-Wei Hai, Yijun Wang, Xiu-Shen WeiICML 2026
- From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language ModelsYuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang 等ACM MM 2025 · 被引用 5 次
- Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language ModelsHairui Ren, Zixuan Wang, Yibo Yang, He Zhao 等ICLR 2026
