MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
Shenyi Zhang, Keyan Guo, Zihao Wang, Xuebin Li, Lingchen Zhao, Hongxin Hu, Chao Shen, Qian Wang
摘要
Multimodal large language models (MLLMs) have demonstrated impressive capabilities across a wide range of generative tasks. However, introducing non-text modalities poses significant challenges for safety alignment. MLLMs often refuse unsafe text-only prompts while producing harmful responses to semantically equivalent multimodal inputs. Existing mitigation strategies, including external guardrails and safety-oriented fine-tuning, largely overlook the mechanisms driving this safety disparity. External guardrails merely circumvent the model's intrinsic defects, while safety-oriented finetuning treats alignment as a black-box optimization problem, failing to diagnose and repair this flaw specifically. As a result, these methods either incur substantial inference latency with limited protection or significantly compromise model utility and rely heavily on large-scale multimodal datasets. Consequently, achieving effective and utility-preserving safety alignment for MLLMs remains an unresolved challenge. In this paper, we conduct a geometric analysis of MLLM representations to investigate the causes of multimodal safety degradation. We reveal that the safety mechanisms learned in the text-only modality actually persist in the multimodal setting. The safety subspace with refusal boundary remains valid across different modalities, where representations falling inside this boundary consistently elicit safe refusal responses. However, we observe a critical shift in representation where most unsafe multimodal inputs fall outside the boundary and bypass this intrinsic mechanism. This finding identifies the root cause of safety failure in the multimodal setting as a representation shift rather than a lack of safety capability. Motivated by this observation, we propose MMAligner, an MLLM safeguarding method based on representation calibration. Unlike existing methods, MMAligner addresses this representation shift by optimizing the model to map unsafe multimodal representations inside the pre-existing refusal boundary. Specifically, it adopts a hard lower bound to ensure refusal and a soft upper bound to prevent excessive modification, while preserving the representation of benign inputs. Extensive experiments * Qian Wang is the corresponding author. on multiple open-source MLLMs demonstrate that MMAligner increases the average refusal rate of multimodal unsafe inputs to 99% while incurring less than 2% utility degradation using minimal data, significantly outperforming existing baselines in the safety-utility trade-off. CCS Concepts • Security and privacy → Software and application security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
相关 Paper
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie 等EMNLP 2024 · 被引用 21 次
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang 等ACL 2025
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 被引用 1 次
- Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMsMing Wen, Kun Yang, Xin Chen, Jingyu Zhang 等ICLR 2026 · 被引用 4 次
- MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and ModalitiesSahil Verma, Keegan Hines, Jeff A. Bilmes, Charlotte Siska 等EMNLP 2025
