VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
Ziyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du, Jinguo Zhu, Han Liu, Jinghui Chen, Ting Wang, Fenglong Ma
摘要
Vision-Language (VL) pre-trained models have shown their superiority on many multimodal tasks. However, the adversarial robustness of such models has not been fully explored. Existing approaches mainly focus on exploring the adversarial robustness under the white-box setting, which is unrealistic. In this paper, we aim to investigate a new yet practical task to craft image and text perturbations using pre-trained VL models to attack black-box fine-tuned models on different downstream tasks. Towards this end, we propose VLAttack to generate adversarial samples by fusing perturbations of images and texts from both single-modal and multimodal levels. At the single-modal level, we propose a new block-wise similarity attack (BSA) strategy to learn image perturbations for disrupting universal representations. Besides, we adopt an existing text attack strategy to generate text perturbations independent of the image-modal attack. At the multimodal level, we design a novel iterative cross-search attack (ICSA) method to update adversarial image-text pairs periodically, starting with the outputs from the single-modal level. We conduct extensive experiments to attack three widely-used VL pretrained models for six tasks on eight datasets. Experimental results show that the proposed VLAttack framework achieves the highest attack success rates on all tasks compared with state-of-the-art baselines, which reveals a significant blind spot in the deployment of pre-trained VL models. Codes will be released soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsYuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun 等NeurIPS 2024 · 被引用 67 次
- Attention! Your Vision Language Model Could Be Maliciously ManipulatedXiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo 等NeurIPS 2025 · 被引用 13 次
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMsSen Nie, Jie Zhang, Jianxin Yan, Shiguang Shan 等CVPR 2026 · 被引用 9 次
- Learning Robust Vision-Language Models from Natural Latent SpacesZhangyun Wang, Ni Ding, Aniket MahantiNeurIPS 2025 · 被引用 3 次
- Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language ModelsYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Towards Adversarial Attack on Vision-Language Pre-training ModelsJiaming Zhang, Qi Yi, Jitao SangACM MM 2022 · 被引用 111 次
- A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training ModelsHaonan Zheng, Xinyang Deng, Wen Jiang, Wenrui LiACM MM 2024 · 被引用 4 次
- VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained ModelsZiyi Yin, Muchao Ye, Tianrong Zhang, Jiaqi Wang 等AAAI 2024 · 被引用 20 次
- Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training ModelsDong Lu, Zhiqiang Wang, Teng Wang, Weili Guan 等ICCV 2023 · 被引用 141 次
- GLEAM: Enhanced Transferable Adversarial Attacks for Vision-Language Pre-Training Models via Global-Local TransformationsYunqi Liu, Xue Ouyang, Xiaohui CuiICCV 2025 · 被引用 9 次
