VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
Hefei Mei, Zirui Wang, Shen You, Minjing Dong, Chang Xu
Abstract
Large Vision-Language Models (LVLMs) have demonstrated capabilities in multimodal understanding, yet their vulnerability to adversarial attacks raises significant concerns. To achieve practical attacking, this paper aims at efficient and transferable untargeted attacks under limited perturbation sizes. Considering this objective, white-box attacks require full-model gradients and task-specific labels, making costs scale with tasks, while black-box attacks rely on proxy models, typically requiring large perturbation sizes and elaborate transfer strategies. Given the centrality and widespread reuse of the vision encoder in LVLMs, we adopt a gray-box setting that targets the vision encoder alone for efficient but effective attacking. We theoretically establish the feasibility of vision-encoder-only attacks, laying the foundation for our gray-box setting. Based on this analysis, we propose perturbing patch tokens rather than the class token, informed by both theoretical and empirical insights. We generate adversarial examples by minimizing the cosine similarity between clean and perturbed visual features, without accessing the subsequent models, tasks, or labels. This significantly reduces computational overhead while eliminating the task and label dependence. VEAttack has achieved a performance degradation of 94.5% on image caption task and 75.7% on visual question answering task. We also reveal some key observations to provide insights into LVLM attack/defense: 1) hidden layer variations of LLM, 2) token attention differential, 3) Möbius band in transfer attack, 4) low sensitivity to attack steps. The code is available at https://github.com/hefeimei06/VEAttack-LVLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdad0002-ae12-46b3-9624-bcf4db58c4d5Cited by top-tier papers6
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentXiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang et al.NeurIPS 2025 · 49 citations
- PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and AttentionHefei Mei, Zirui Wang, Chang Xu, Jianyuan Guo et al.CVPR 2026 · 3 citations
- On the Adversarial Robustness of Large Vision-Language Models under Visual Token CompressionXinwei Zhang, Hangcheng Liu, Li Bai, Hao Wang et al.ICML 2026 · 2 citations
- SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal ModelsAafiya Shamshad Hussain, Gaurav Srivastava, Alvi Md. Ishmam, Zaber Ibn Abdul Hakim et al.ACL 2026 · 1 citation
- Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention HijackingJingru Li, Wei Ren, Tianqing ZhuACL 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language ModelsYubo Wang, Chaohu Liu, Yanqiu Qu, Haoyu Cao et al.ACM MM 2024 · 10 citations
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsXiaowen Cai, Daizong Liu, Xiaoye Qu, Xiang Fang et al.NeurIPS 2025 · 8 citations
- On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial AttackZhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu et al.USENIX Security 2026
- Attacking Gray-Box Large Vision-Language Models with Adaptive SVD-Structured Adversarial AlignmentDaizong Liu, Xiaowen Cai, Junhao Dong, Zhongliang Guo et al.ICML 2026
- Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language ModelsDaizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou et al.NeurIPS 2024 · 51 citations
