Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
Han Wang, Gang Wang, Huan Zhang
Abstract
Vision Language Models (VLMs) can produce unintended and harmful content when exposed to adversarial attacks, particularly because their vision capabilities create new vulnerabilities. Existing defenses, such as input preprocessing, adversarial training, and response evaluation-based methods, are often impractical for real-world deployment due to their high costs. To address this challenge, we propose ASTRA, an efficient and effective defense by adaptively steering models away from adversarial feature directions to resist VLM attacks. Our key procedures involve finding transferable steering vectors representing the direction of harmful response and applying adaptive activation steering to remove these directions at inference time. To create effective steering vectors, we randomly ablate the visual tokens from the adversarial images and identify those most strongly associated with jailbreaks. These tokens are then used to construct steering vectors. During inference, we perform the adaptive steering method that involves the projection between the steering vectors and calibrated activation, resulting in little performance drops on benign inputs while strongly avoiding harmful outputs under adversarial inputs. Extensive experiments across multiple models and baselines demonstrate our state-of-theart performance and high efficiency in mitigating jailbreak risks. Additionally, ASTRA exhibits good transferability, defending against unseen attacks (i.e., structured-based attack, perturbation-based attack with project gradient descent variants, and text-only attack). Our code is available at https://github.com/ASTRAL-Group/ASTRA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f21cb886-12e8-4363-9b17-71e2ae5307f5Cited by top-tier papers22
- AlphaSteer: Learning Refusal Steering with Principled Null-Space ConstraintLeheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang et al.ICLR 2026 · 52 citations
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt et al.ICLR 2026 · 28 citations
- Understanding and Rectifying Safety Perception Distortion in VLMsXiaohan Zou, Jian Kang, George Kesidis, Lu LinNeurIPS 2025 · 20 citations
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsJingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui et al.NeurIPS 2025 · 17 citations
- Learning to Steer: Input-dependent Steering for Multimodal LLMsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny et al.NeurIPS 2025 · 14 citations
Builds on26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language ModelsXingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang et al.CVPR 2026 · 5 citations
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak DefenderWeixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng et al.EMNLP 2025
- Activation Manipulation Attack: Penetrating and Harmful Jailbreak Attack Against Large Vision-Language ModelsHaojie Hao, Jiakai Wang, Aishan Liu, Yuqing Ma et al.AAAI 2026
- InDe-LLM: Defending against Jailbreak Attacks in LLM-Powered Systems via Intention DisentanglingYujue Wang, Quan Zhang, Chijin Zhou, Gwihwan Go et al.FSE 2026
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang et al.AAAI 2026
