VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained Models
Ziyi Yin, Muchao Ye, Tianrong Zhang, Jiaqi Wang, Han Liu, Jinghui Chen, Ting Wang, Fenglong Ma
Abstract
Visual Question Answering (VQA) is a fundamental task in computer vision and natural language process fields. Although the “pre-training & finetuning” learning paradigm significantly improves the VQA performance, the adversarial robustness of such a learning paradigm has not been explored. In this paper, we delve into a new problem: using a pre-trained multimodal source model to create adversarial image-text pairs and then transferring them to attack the target VQA models. Correspondingly, we propose a novel VQATTACK model, which can iteratively generate both im- age and text perturbations with the designed modules: the large language model (LLM)-enhanced image attack and the cross-modal joint attack module. At each iteration, the LLM-enhanced image attack module first optimizes the latent representation-based loss to generate feature-level image perturbations. Then it incorporates an LLM to further enhance the image perturbations by optimizing the designed masked answer anti-recovery loss. The cross-modal joint attack module will be triggered at a specific iteration, which updates the image and text perturbations sequentially. Notably, the text perturbation updates are based on both the learned gradients in the word embedding space and word synonym-based substitution. Experimental results on two VQA datasets with five validated models demonstrate the effectiveness of the proposed VQATTACK in the transferable attack setting, compared with state-of-the-art baselines. This work reveals a significant blind spot in the “pre-training & fine-tuning” paradigm on VQA tasks. The source code can be found in the link https://github.com/ericyinyzy/VQAttack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language ModelJinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan et al.NeurIPS 2025 · 3 citations
- The Power of Decaying Steps: Enhancing Attack Stability and Transferability for Sign-based OptimizersWei Tao, Yang Dai, Jincai Huang, Qing TaoCVPR 2026
- From Zero to Hero: Cross-modal-enhanced Adversarial Item Promotion Attack against Multimodal Recommender SystemsMengyu Yao, Ziqi Zhang, Yifeng Cai, Junlin Liu et al.USENIX Security 2026
- HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained ModelsHan Liu, Jiaqi Li, Zhi Xu, Xiaotong Zhang et al.NeurIPS 2025
- Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language ModelsHao Cheng, Erjia Xiao, Jiayan Yang, Jinhao Duan et al.ACM MM 2025
Builds on19
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu et al.NeurIPS 2022 · 790 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
Related papers
- VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained ModelsZiyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du et al.NeurIPS 2023 · 109 citations
- Towards Adversarial Attack on Vision-Language Pre-training ModelsJiaming Zhang, Qi Yi, Jitao SangACM MM 2022 · 111 citations
- GLEAM: Enhanced Transferable Adversarial Attacks for Vision-Language Pre-Training Models via Global-Local TransformationsYunqi Liu, Xue Ouyang, Xiaohui CuiICCV 2025 · 9 citations
- A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training ModelsHaonan Zheng, Xinyang Deng, Wen Jiang, Wenrui LiACM MM 2024 · 4 citations
- Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal OptimizationXiang Fang, Wanlong Fang, Changshuo WangAAAI 2026 · 3 citations
