ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
Weikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li, Huiping Zhuang, Ruidong Wang, Cen Chen, Hao Peng
Abstract
Multimodal Large Language Models (MLLMs) are increasingly vulnerable to multimodal Indirect Prompt Injection (IPI) attacks, which embed malicious instructions in images, videos, or audio to hijack model behavior. Existing defenses, designed primarily for text-only LLMs, are unsuitable for countering these multimodal threats, as they are easily bypassed, modality-dependent, or generalize poorly. Inspired by activation steering researches, we hypothesize that a robust, general defense independent of modality can be achieved by steering the model's behavior in the representation space. Through extensive experiments, we discover that the instruction-following behavior of MLLMs is encoded in a subspace. Steering along directions within this subspace can enforce adherence to user instructions, forming the basis of a defense. However, we also found that a naive defense direction could be coupled with a utility-degrading direction, and excessive intervention strength harms model performance. To address this, we propose ARGUS, which searches for an optimal defense direction within the safety subspace that decouples from the utility degradation direction, further combining adaptive strength steering to achieve a better safety-utility trade-off. ARGUS also introduces lightweight injection detection stage to to activate the defense on-demand, and a post-filtering stage to verify defense success. Experimental results show that ARGUS can achieve robust defense against multimodal IPI while maximally preserving the MLLM's utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd9666cb-5e1b-45eb-85f8-b102ef582351Builds on22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- EEGPT: Pretrained Transformer for Universal and Reliable Representation of EEG SignalsGuangyu Wang, Wenchao Liu, Yuhong He, Cong Xu et al.NeurIPS 2024 · 267 citations
- On the Exploitability of Instruction TuningManli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping et al.NeurIPS 2023 · 166 citations
- VACE: All-in-One Video Creation and EditingZeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang et al.ICCV 2025 · 58 citations
Related papers
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMsYinan Zhong, Qianhao Miao, Yanjiao Chen, Jiangyi Deng et al.NDSS 2026 · 13 citations
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie et al.EMNLP 2024 · 21 citations
- Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language ModelsXingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang et al.CVPR 2026 · 5 citations
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He et al.ACL 2025
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou et al.EMNLP 2025
