Defending LVLMs Against Vision Attacks Through Partial-Perception Supervision
Qi Zhou, Dongxia Wang, Tianlin Li, Yun Lin, Yang Liu, Jin Song Dong, Qing Guo
Abstract
Recent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especially cropping, using majority voting across responses of modified images as corrected responses. However, these modifications often result in partial images and distort the semantics, which reduces response quality on clean images after voting. Instead of directly using responses from partial images for voting, we investigate using them to supervise (guide) the LVLM's responses to the original images at inference time. We propose a blackbox, training-free method called DPS (Defense through Partial-Perception Supervision). In this approach, the model is prompted using the responses generated by a model that perceives only a partial image. With DPS, the model can adjust its response based on partial image understanding when under attack, while confidently maintaining its original response for clean input. Empirical experiments show our method outperforms the baseline, cutting the average attack success rate by 76.3% across six datasets on three popular models. Our code is available at https: //github.com/tools-only/DPS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsXiaowen Cai, Daizong Liu, Xiaoye Qu, Xiang Fang et al.NeurIPS 2025 · 8 citations
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksShide Zhou, Tianlin Li, Kailong Wang, Yihao Huang et al.ICSE 2025 · 3 citations
- On the Adversarial Robustness of Large Vision-Language Models under Visual Token CompressionXinwei Zhang, Hangcheng Liu, Li Bai, Hao Wang et al.ICML 2026 · 2 citations
- SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World EnvironmentsYue Cao, Yun Xing, Jie Zhang, Di Lin et al.CVPR 2025
Builds on19
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language ModelsHefei Mei, Zirui Wang, Shen You, Minjing Dong et al.ICLR 2026 · 9 citations
- One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention HeadJunhao Xia, Haotian Zhu, Shuchao Pang, Zhigang Lu et al.NeurIPS 2025 · 5 citations
- Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language ModelsYubo Wang, Chaohu Liu, Yanqiu Qu, Haoyu Cao et al.ACM MM 2024 · 10 citations
- Attacking Gray-Box Large Vision-Language Models with Adaptive SVD-Structured Adversarial AlignmentDaizong Liu, Xiaowen Cai, Junhao Dong, Zhongliang Guo et al.ICML 2026
- Transferable Direct Prompt Injection via Activation-Guided MCMC SamplingMinghui Li, Hao Zhang, Yechao Zhang, Wei Wan et al.EMNLP 2025 · 1 citation
