Pro-Cap: Leveraging a Frozen Vision-Language Model for Hateful Meme Detection
Rui Cao, Ming Shan Hee, Adriel Kuek, Wen-Haw Chong, Roy Ka-Wei Lee, Jing Jiang
Abstract
Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to fine-tune pre-trained vision-language models (PVLMs) for this task. However, with increasing model sizes, it becomes important to leverage powerful PVLMs more efficiently, rather than simply fine-tuning them. Recently, researchers have attempted to convert meme images into textual captions and prompt language models for predictions. This approach has shown good performance but suffers from noninformative image captions. Considering the two factors mentioned above, we propose a probing-based captioning approach to leverage PVLMs in a zero-shot visual question answering (VQA) manner. Specifically, we prompt a frozen PVLM by asking hateful contentrelated questions and use the answers as image captions (which we call Pro-Cap), so that the captions contain information critical for hateful content detection. The good performance of models with Pro-Cap on three benchmarks validates the effectiveness and generalization of the proposed method. 1
• Computing methodologies → Natural language processing; Computer vision representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a72e2fff-229a-4938-bc79-61ae97283aebCited by top-tier papers24
- Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language ModelsHongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma et al.WWW 2024 · 43 citations
- MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and BilibiliHan Wang, Tan Rui Yang, Usman Naseem, Roy Ka-Wei LeeACM MM 2024 · 23 citations
- MIND: A Multi-agent Framework for Zero-shot Harmful Meme DetectionZiyan Liu, Chunxiao Fan, Haoran Lou, Yuexin Wu et al.ACL 2025 · 15 citations
- Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate DetectionJian Lang, Rongpei Hong, Jin Xu, Yili Li et al.WWW 2025 · 14 citations
- Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in MemesXuanrui Lin, Chao Jia, Junhui Ji, Hui Han et al.WWW 2025 · 9 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Prompting for Multimodal Hateful Meme ClassificationRui Cao, Roy Ka-Wei Lee, Wen-Haw Chong, Jing JiangEMNLP 2022 · 68 citations
- From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language ModelsYihan Ma, Xinyue Shen, Yiting Qu, Ning Yu et al.USENIX Security 2025
- Transferable Decoding with Visual Entities for Zero-Shot Image CaptioningJunjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He et al.ICCV 2023 · 80 citations
- PromptMTopic: Unsupervised Multimodal Topic Modeling of Memes using Large Language ModelsNirmalendu Prakash, Han Wang, Nguyen-Khoi Hoang, Ming Shan Hee et al.ACM MM 2023 · 24 citations
- A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language ModelsWoojeong Jin, Yu Cheng, Yelong Shen, Weizhu Chen et al.ACL 2022
