Medical Vision-Language Pre-training with Multimodal Variational Masked Autoencoder for Robust Medical VQA
Dexuan Xu, Yanyuan Chen, Yu Huang, Shihao E, Yiwei Lou, Yongzhi Cao, Hanpin Wang, Meikang Qiu
Abstract
Medical Visual Question Answering (Medical VQA) plays an important role in medical informatics. However, the robustness of existing medical VQA models is severely challenged by adversarial attacks. Current methods (e.g. adversarial training and noise-based reasoning) heavily rely on additional data or complex procedures and often ignore model-level robustness. To address these issues, we propose Multimodal Variational Masked Autoencoder (MVMAE), a novel pre-training framework designed to enhance the robustness of the medical VQA task. MVMAE leverages masked modeling and variational inference to extract robust multimodal features. The framework introduces a low-cost multimodal bottleneck fusion module and employs reparameterization to sample robust latent representations, ensuring effective feature fusion and reconstruction. Extensive experiments on public medical VQA datasets demonstrate that MVMAE significantly improves resistance to various adversarial attacks and outperforms other medical multimodal pre-training methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f7bd8b8d-a443-4aaa-8bdc-d977b07edc46Related papers
- VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained ModelsZiyi Yin, Muchao Ye, Tianrong Zhang, Jiaqi Wang et al.AAAI 2024 · 20 citations
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question AnsweringYuanhao Zou, Zhaozheng YinCVPR 2025
- Improving VAEs' Robustness to Adversarial AttackMatthew Willetts, Alexander Camuto, Tom Rainforth, Stephen J. Roberts et al.ICLR 2021 · 30 citations
- Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA ModelsLinjie Li, Jie Lei, Zhe Gan, Jingjing LiuICCV 2021 · 99 citations
- On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial AttackZhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu et al.USENIX Security 2026
