Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
DongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo Yu
Abstract
Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet most evaluations rely on artificial images. This study asks: How safe are current VLMs when confronted with meme images that ordinary users share? To investigate this question, we introduce MEMESAFETYBENCH, a 50,430-instance benchmark pairing real meme images with both harmful and benign instructions. Using a comprehensive safety taxonomy and LLMbased instruction generation, we assess multiple VLMs across single and multi-turn interactions. We investigate how real-world memes influence harmful outputs, the mitigating effects of conversational context, and the relationship between model scale and safety metrics. Our findings demonstrate that VLMs are more vulnerable to meme-based harmful prompts than to synthetic or typographic images. Memes significantly increase harmful responses and decrease refusals compared to textonly inputs. Though multi-turn interactions provide partial mitigation, elevated vulnerability persists. These results highlight the need for ecologically valid evaluations and stronger safety mechanisms. MEMESAFETYBENCH is publicly available at https://github.com/ oneonlee/Meme-Safety-Bench . Warning: This paper includes examples of harmful language and images that may be sensitive or uncomfortable. Reader discretion is recommended. Meme (๐ผ๐ผ ๐๐ ) False or Misleading Information ๐๐ ๐๐ Category (๐๐ ๐๐ ) Classification GPT-4o โ Harmful/Harmless Instruction Generation & Verification โก Task (๐ก๐ก ๐๐ ๐๐
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f266950-db6d-4778-925f-fabfcfbe0e4fCited by top-tier papers2
- Jailbreaking on Text-to-Video Models via Scene Splitting StrategyWonjun Lee, Haon Park, Doehyeon Lee, Bumsub Ham et al.ICLR 2026 ยท 10 citations
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMsDasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt et al.ACL 2026 ยท 3 citations
Builds on19
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 ยท 1,031 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 ยท 1,022 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 ยท 1,016 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 ยท 722 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 ยท 404 citations
Related papers
- From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language ModelsYihan Ma, Xinyue Shen, Yiting Qu, Ning Yu et al.USENIX Security 2025
- MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language ModelsZhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao et al.EMNLP 2025
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang et al.ACL 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyWonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu et al.ICML 2025
