MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
Shenyi Zhang, Keyan Guo, Zihao Wang, Xuebin Li, Lingchen Zhao, Hongxin Hu, Chao Shen, Qian Wang
Abstract
Multimodal large language models (MLLMs) have demonstrated impressive capabilities across a wide range of generative tasks. However, introducing non-text modalities poses significant challenges for safety alignment. MLLMs often refuse unsafe text-only prompts while producing harmful responses to semantically equivalent multimodal inputs. Existing mitigation strategies, including external guardrails and safety-oriented fine-tuning, largely overlook the mechanisms driving this safety disparity. External guardrails merely circumvent the model's intrinsic defects, while safety-oriented finetuning treats alignment as a black-box optimization problem, failing to diagnose and repair this flaw specifically. As a result, these methods either incur substantial inference latency with limited protection or significantly compromise model utility and rely heavily on large-scale multimodal datasets. Consequently, achieving effective and utility-preserving safety alignment for MLLMs remains an unresolved challenge. In this paper, we conduct a geometric analysis of MLLM representations to investigate the causes of multimodal safety degradation. We reveal that the safety mechanisms learned in the text-only modality actually persist in the multimodal setting. The safety subspace with refusal boundary remains valid across different modalities, where representations falling inside this boundary consistently elicit safe refusal responses. However, we observe a critical shift in representation where most unsafe multimodal inputs fall outside the boundary and bypass this intrinsic mechanism. This finding identifies the root cause of safety failure in the multimodal setting as a representation shift rather than a lack of safety capability. Motivated by this observation, we propose MMAligner, an MLLM safeguarding method based on representation calibration. Unlike existing methods, MMAligner addresses this representation shift by optimizing the model to map unsafe multimodal representations inside the pre-existing refusal boundary. Specifically, it adopts a hard lower bound to ensure refusal and a soft upper bound to prevent excessive modification, while preserving the representation of benign inputs. Extensive experiments * Qian Wang is the corresponding author. on multiple open-source MLLMs demonstrate that MMAligner increases the average refusal rate of multimodal unsafe inputs to 99% while incurring less than 2% utility degradation using minimal data, significantly outperforming existing baselines in the safety-utility trade-off. CCS Concepts • Security and privacy → Software and application security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on35
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
Related papers
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie et al.EMNLP 2024 · 21 citations
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang et al.ACL 2025
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 1 citation
- Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMsMing Wen, Kun Yang, Xin Chen, Jingyu Zhang et al.ICLR 2026 · 4 citations
- MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and ModalitiesSahil Verma, Keegan Hines, Jeff A. Bilmes, Charlotte Siska et al.EMNLP 2025
