SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models
Yuxin Gou, Xiaoning Dong, Qin Li, Shishen Gu, Richang Hong, Wenbo Hu
Abstract
Multimodal large language models (MLLMs) demonstrate impressive capabilities by integrating visual and textual information. However, the incorporation of visual modalities also introduces new and complex safety risks, rendering even the most advanced models vulnerable to sophisticated jailbreak attacks. This paper first analyzes the impact of inserting safety reasoning prompt on various aspects of the model. We find that this external method can help the model resist jailbreak attacks to some extent, but the model still fails to distinguish specific semantic scenarios, resulting in a significantly increased refusal rate for benign queries. Inspired by this, we propose a novel training framework, SURE (Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models), designed to help models internalize chain-of-thought-based safety decision-making capabilities. Extensive experiments demonstrate that SURE significantly improves model safety while effectively avoiding over-defense, achieving a good balance between safety and generality. Finally, we create a large-scale multimodal safety reasoning dataset, MLLM-SCoT-Plus, to facilitate research on safety alignment in multimodal models. Our code and the dataset are publicly available at https://github.com/hfutml/SURE . Warning: This paper contains offensive and harmful examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27a02cf8-7e04-4683-8c41-fb14ebd8bed9Cited by top-tier papers4
- Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective MemoryCe Zhang, Jinxi He, Junyi He, Katia Sycara et al.CVPR 2026 · 5 citations
- ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text ReasoningLitao Guo, Jinsong Zhou, Shuaibo Li, Man CHEN et al.ICML 2026
- ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMsZijun Chen, Wenbo Hu, Ya Li, Lei Miao et al.ICLR 2026
- Mitigating Safety Context Amnesia in Multimodal Reasoning Models via Intent-Guided Safety ReasoningXiyao Dong, Guangsheng Cheng, YiLong Chen, Xiaojin Zhang et al.ACL 2026
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Image Hijacks: Adversarial Images can Control Generative Models at RuntimeLuke Bailey, Euan Ong, Stuart Russell, Scott EmmonsICML 2024 · 171 citations
- Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLPYangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi et al.EMNLP 2022 · 28 citations
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie et al.EMNLP 2024 · 21 citations
Related papers
- Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentSoumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan et al.CVPR 2025
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore MechanismBeitao Chen, Xinyu Lyu, Shengming Yuan, Jingkuan Song et al.NeurIPS 2025 · 14 citations
- Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising UtilityMengxuan Wang, Yuxin Chen, Gang Xu, Tao He et al.ICML 2026
- Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMsMing Wen, Kun Yang, Xin Chen, Jingyu Zhang et al.ICLR 2026 · 4 citations
- TVChain: Leveraging Textual-Visual Prompt Chains for Jailbreaking Large Vision-Language ModelsHao Yu, Ke Liang, Junxian Duan, Jun Wang et al.AAAI 2026
