Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
Xingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang, Kesen Zhao, Hanwang Zhang, Xiangnan He
Abstract
As vision-language models (VLMs) are increasingly deployed in open-world scenarios, they can be easily induced by visual jailbreak attacks to generate harmful content, posing serious risks to model safety and trustworthy usage. Recent activation steering methods inject directional vectors into model activations during inference to induce refusal behaviors and have demonstrated effectiveness. However, a steering vector may both enhance refusal ability and cause over-refusal, thereby degrading model performance on benign inputs. Moreover, due to the lack of theoretical interpretability, these methods still suffer from limited robustness and effectiveness. To better balance safety and utility, we propose NullSteer, a null-space projected activation defense framework. Our method constructs refusal directions within model activations through a linear transformation: it maintains zero perturbation within the benign subspace while dynamically inducing refusal along potentially harmful directions, thereby theoretically achieving safety enhancement without impairing the model's general capabilities. Extensive experiments show that NullSteer significantly reduces harmful outputs under various jailbreak attacks (average ASR reduction over 15 percent on MiniGPT-4) while maintaining comparable performance to the original model on general benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76995b81-c162-4f0f-8a09-3fade9f65ff9Cited by top-tier papers2
- Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM UnlearningNaixin Zhai, Pengyang Shao, Binbin Zheng, Yonghui Yang et al.ACL 2026 · 10 citations
- Robustifying Vision-Language Models via Test-Time Prompt AdaptationXingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu et al.ICML 2026 · 1 citation
Builds on32
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against JailbreaksHan Wang, Gang Wang, Huan ZhangCVPR 2025
- AlphaSteer: Learning Refusal Steering with Principled Null-Space ConstraintLeheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang et al.ICLR 2026 · 52 citations
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak DefenderWeixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng et al.EMNLP 2025
- Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation SimulationLUOYU CHEN, Weiqi Wang, Zhiyi Tian, Chenhan Zhang et al.ICML 2026
- Jailbreaking the Matrix: Nullspace Steering for Controlled Model SubversionVishal Pramanik, Maisha Maliha, Susmit Jha, Sumit Kumar JhaICLR 2026 · 3 citations
