Automating Steering for Safe Multimodal Large Language Models
Lyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng
Abstract
Recent progress in Multimodal Large Language Models (MLLMs) has unlocked powerful cross-modal reasoning abilities, but also raised new safety concerns, particularly when faced with adversarial multimodal inputs. To improve the safety of MLLMs during inference, we introduce a modular and adaptive inferencetime intervention technology, AutoSteer, without requiring any fine-tuning of the underlying model. AutoSteer incorporates three core components: (1) a novel Safety Awareness Score (SAS) that automatically identifies the most safety-relevant distinctions among the model's internal layers; (2) an adaptive safety prober trained to estimate the likelihood of toxic outputs from intermediate representations; and (3) a lightweight Refusal Head that selectively intervenes to modulate generation when safety risks are detected. Experiments on LLaVA-OV and Chameleon across diverse safety-critical benchmarks demonstrate that AutoSteer significantly reduces the Attack Success Rate (ASR) for textual, visual, and cross-modal threats, while maintaining general abilities. These findings position AutoSteer as a practical, interpretable, and effective framework for safer deployment of multimodal AI systems 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36336931-ead5-4ce1-82d2-85200f5a3b0dCited by top-tier papers5
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use AgentsTri Cao, Bennett Lim, Yue Liu, Yuan Sui et al.ICLR 2026 · 45 citations
- Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language ModelsXingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang et al.CVPR 2026 · 5 citations
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language ModelsHongji Li, Manjiang Yu, Junchi Yao, PRIYANKA SINGH et al.CVPR 2026 · 3 citations
- Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising UtilityMengxuan Wang, Yuxin Chen, Gang Xu, Tao He et al.ICML 2026
- How Controllable Are Large Language Models? A Unified Evaluation across Behavioral GranularitiesZiwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong et al.ACL 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsShaona Ghosh, Amrita Bhattacharjee, Yftah Ziser, Christopher ParisienEMNLP 2025
- SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal SteeringWeilin Lin, Jianze Li, Hui Xiong, Li LiuICML 2026 · 6 citations
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 1 citation
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li et al.ACL 2026
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie et al.EMNLP 2024 · 21 citations
