Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, Zhaopeng Tu
摘要
Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a significant challenge, particularly in accurately identifying whether multimodal content is safe or unsafe-a capability we term safety awareness. In this paper, we introduce MMSafeAware, the first comprehensive multimodal safety awareness benchmark designed to evaluate MLLMs across 29 safety scenarios with 1,500 carefully curated image-prompt pairs. MMSafeAware includes both unsafe and over-safety subsets to assess models' abilities to correctly identify unsafe content and avoid over-sensitivity that can hinder helpfulness. Evaluating nine widely used MLLMs using MMSafeAware reveals that current models are not sufficiently safe and often overly sensitive; for example, GPT-4V misclassifies 36.1% of unsafe inputs as safe and 59.9% of benign inputs as unsafe. We further explore three methods to improve safety awareness-prompting-based approaches, visual contrastive decoding, and vision-centric reasoning fine-tuning-but find that none achieve satisfactory performance. Our findings highlight the profound challenges in developing MLLMs with robust safety awareness, underscoring the need for further research in this area. All the code and data is publicly available 1 to facilitate future research. WARNING: This paper contains unsafe contents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective MemoryCe Zhang, Jinxi He, Junyi He, Katia Sycara 等CVPR 2026 · 被引用 5 次
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language ModelsGuolei Huang, Qinzhi Peng, Gan Xu, Yao Huang 等CVPR 2026 · 被引用 2 次
- Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMsJingyu Zhang, Kun Yang, Ming Wen, Zhuoer Xu 等ICLR 2026
- HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden StatesYilei Jiang, Xinyan Gao, Tianshuo Peng, Yingshui Tan 等ACL 2025
它引用的顶会 Paper7
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David BammanEMNLP 2023 · 被引用 70 次
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu 等FSE 2023 · 被引用 50 次
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li 等ICCV 2025 · 被引用 37 次
- SafeText: A Benchmark for Exploring Physical Safety in Language ModelsSharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton 等EMNLP 2022 · 被引用 14 次
相关 Paper
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang 等ACL 2025
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas 等ICLR 2025
- The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image ReasoningRenmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang 等ACL 2026 · 被引用 1 次
- Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsYikang Zhou, Tao Zhang, Shilin Xu, Shihao Chen 等ICCV 2025 · 被引用 2 次
- When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation ParadigmYe Leng, Junjie Chu, Mingjie Li, Chenhao Lin 等CVPR 2026 · 被引用 3 次
