Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
DongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo Yu
摘要
Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet most evaluations rely on artificial images. This study asks: How safe are current VLMs when confronted with meme images that ordinary users share? To investigate this question, we introduce MEMESAFETYBENCH, a 50,430-instance benchmark pairing real meme images with both harmful and benign instructions. Using a comprehensive safety taxonomy and LLMbased instruction generation, we assess multiple VLMs across single and multi-turn interactions. We investigate how real-world memes influence harmful outputs, the mitigating effects of conversational context, and the relationship between model scale and safety metrics. Our findings demonstrate that VLMs are more vulnerable to meme-based harmful prompts than to synthetic or typographic images. Memes significantly increase harmful responses and decrease refusals compared to textonly inputs. Though multi-turn interactions provide partial mitigation, elevated vulnerability persists. These results highlight the need for ecologically valid evaluations and stronger safety mechanisms. MEMESAFETYBENCH is publicly available at https://github.com/ oneonlee/Meme-Safety-Bench . Warning: This paper includes examples of harmful language and images that may be sensitive or uncomfortable. Reader discretion is recommended. Meme (𝐼𝐼 𝑖𝑖 ) False or Misleading Information 𝑐𝑐 𝑖𝑖 Category (𝑐𝑐 𝑖𝑖 ) Classification GPT-4o ① Harmful/Harmless Instruction Generation & Verification ② Task (𝑡𝑡 𝑖𝑖 𝑗𝑗
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Jailbreaking on Text-to-Video Models via Scene Splitting StrategyWonjun Lee, Haon Park, Doehyeon Lee, Bumsub Ham 等ICLR 2026 · 被引用 10 次
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMsDasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper19
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang 等NeurIPS 2023 · 被引用 404 次
相关 Paper
- From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language ModelsYihan Ma, Xinyue Shen, Yiting Qu, Ning Yu 等USENIX Security 2025
- MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language ModelsZhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao 等EMNLP 2025
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang 等ACL 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun 等ACL 2024
- ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyWonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu 等ICML 2025
