Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language Models
Daizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou, Xiang Fang, Keke Tang, Yao Wan, Lichao Sun
Abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding tasks. Nevertheless, these models are susceptible to adversarial examples. In real-world applications, existing LVLM attackers generally rely on the detailed prior knowledge of the model to generate effective perturbations. Moreover, these attacks are task-specific, leading to significant costs for designing perturbation. Motivated by the research gap and practical demands, in this paper, we make the first attempt to build a universal attacker against real-world LVLMs, focusing on two critical aspects: ( i ) restricting access to only the LVLM inputs and outputs. ( ii ) devising a universal adversarial patch, which is task-agnostic and can deceive any LVLM-driven task when applied to various inputs. Specifically, we start by initializing the location and the pattern of the adversarial patch through random sampling, guided by the semantic distance between their output and the target label. Subsequently, we maintain a consistent patch location while refining the pattern to enhance semantic resemblance to the target. In particular, our approach incorporates a diverse set of LVLM task inputs as query samples to approximate the patch gradient, capitalizing on the importance of distinct inputs. In this way, the optimized patch is universally adversarial against different tasks and prompts, leveraging solely gradient estimates queried from the model. Extensive experiments are conducted to verify the strong universal adversarial capabilities of our proposed attack with prevalent LVLMs including
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83501212-3cf2-4d8e-9b33-ff405b28034cCited by top-tier papers14
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsHai Yan, Haijian Ma, Xiaowen Cai, Daizong Liu et al.NeurIPS 2025 · 21 citations
- Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World TrustworthinessXiang Fang, Wanlong Fang, Wei JiICML 2026 · 17 citations
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 17 citations
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsXiaowen Cai, Daizong Liu, Xiaoye Qu, Xiang Fang et al.NeurIPS 2025 · 8 citations
- Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural NetworksYi Yu, Qixin Zhang, Shuhan Ye, Xun Lin et al.ICLR 2026 · 8 citations
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Spatial-Spectral Homogeneous Attacks on Physical-World Large Vision-Language ModelsDaizong Liu, Baoquan Chen, Wei HuAAAI 2026
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language ModelsHefei Mei, Zirui Wang, Shen You, Minjing Dong et al.ICLR 2026 · 9 citations
- Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal OptimizationXiang Fang, Wanlong Fang, Changshuo WangAAAI 2026 · 3 citations
- Attacking Gray-Box Large Vision-Language Models with Adaptive SVD-Structured Adversarial AlignmentDaizong Liu, Xiaowen Cai, Junhao Dong, Zhongliang Guo et al.ICML 2026
- LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained ModelsAlvi Md. Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Chris ThomasAAAI 2026
