USENIX Security2025Top-tier venue
From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Models
Yihan Ma, Xinyue Shen, Yiting Qu, Ning Yu, Michael Backes, Savvas Zannettou, Yang Zhang
Abstract
Open-source Vision Language Models (VLMs) have rapidly advanced, blending natural language with visual modalities, leading them to achieve remarkable performance on tasks such as image captioning and visual question answering. However, their effectiveness in real-world scenarios remains uncertain, as real-world images-particularly hateful memes-often convey complex semantics, cultural references, and emotional signals far beyond those in experimental datasets. In this paper, we present an in-depth evaluation of VLMs' ability to interpret hateful memes by curating a dataset of 39 hateful memes and 12,775 responses from seven representative VLMs using carefully designed prompts. Our manual annotations of the responses' informativeness and soundness reveal that VLMs can identify visual concepts and understand cultural and emotional backgrounds, especially for the well-known hateful memes. However, we find that the VLMs lack robust safeguards to effectively detect and reject hateful content, making them vulnerable to misuse for generating harmful outputs such as hate speech and offensive slogans. Our findings show that 40% of VLM-generated hate speech and over 10% of hateful jokes and slogans were flagged as harmful, emphasizing the urgent need for stronger safety measures and ethical guidelines to mitigate misuse. We hope our study serves as a foundation for improving VLM safety and ethical standards in handling hateful content. 1 Disclaimer. This paper includes examples of hateful content, including antisemitic symbols and other forms of highly offensive material. Reader discretion is advised when reviewing this content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ab8f013-4ae7-4b51-99c6-0033db952ea1Cited by top-tier papers4
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful IllusionsYiting Qu, Ziqing Yang, Yihan Ma, Michael Backes et al.ICCV 2025 · 6 citations
- SAGE: Synergistic Adaptive Gating of Experts for Hateful Video DetectionJie Huang, Xin Liao, Junjie Wang, Mingyang Li et al.ACL 2026
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across ModalitiesYiting Qu, Michael Backes, Yang ZhangUSENIX Security 2025
- GPTracker: A Large-Scale Measurement of Misused GPTsXinyue Shen, Yun Shen, Michael Backes, Yang ZhangS&P 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 1 citation
- Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme DetectionJingbiao Mei, Jinghong Chen, Guangyu Yang, Weizhe Lin et al.EMNLP 2025 · 2 citations
- Pro-Cap: Leveraging a Frozen Vision-Language Model for Hateful Meme DetectionRui Cao, Ming Shan Hee, Adriel Kuek, Wen-Haw Chong et al.ACM MM 2023 · 54 citations
- PunMemeCN: A Benchmark to Explore Vision-Language Models' Understanding of Chinese Pun MemesZhijun Xu, Siyu Yuan, Yiqiao Zhang, Jingyu Sun et al.EMNLP 2025
- MetaGPT: A Large Vision-Language Model for Meme Metaphor UnderstandingBo Xu, Chenyuan Wang, Xinyu Chen, Hongfei Lin et al.AAAI 2026
