MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
Prince Jha, Raghav Jain, Konika Mandal, Aman Chadha, Sriparna Saha, Pushpak Bhattacharyya
Abstract
In the digital world, memes present a unique challenge for content moderation due to their potential to spread harmful content. Although detection methods have improved, proactive solutions such as intervention are still limited, with current research focusing mostly on textbased content, neglecting the widespread influence of multimodal content like memes. Addressing this gap, we present MemeGuard, a comprehensive framework leveraging Large Language Models (LLMs) and Visual Language Models (VLMs) for meme intervention. MemeGuard harnesses a specially fine-tuned VLM, VLMeme, for meme interpretation, and a multimodal knowledge selection and ranking mechanism (MKS) for distilling relevant knowledge. This knowledge is then employed by a general-purpose LLM to generate contextually appropriate interventions. Another key contribution of this work is the Intervening Cyberbullying in Multimodal Memes (ICMM) dataset, a high-quality, labeled dataset featuring toxic memes and their corresponding humanannotated interventions. We leverage ICMM to test MemeGuard, demonstrating its proficiency in generating relevant and effective responses to toxic memes. 1 Disclaimer: This paper contains harmful content that may be disturbing to some readers. * Work does not relate to position at Amazon. 1 Code and dataset are available at https://github.com/ Jhaprince/MemeGuard USER Ideal Intervention Making fun of someone's intelligence is hurtful and unnecessary. We should strive to treat others with kindness and empathy, rather than belittling them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98282493-f8be-44b8-b8f9-7276941eadf7Cited by top-tier papers7
- ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content ModerationMengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu et al.AAAI 2025 · 14 citations
- MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion RecognitionJian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang et al.ACM MM 2025 · 5 citations
- Toxicity Begets Toxicity: Unraveling Conversational Chains in Political PodcastsNaquee Rizwan, Nayandeep Deb, Sarthak Roy, Vishwajeet Singh Solanki et al.ACM MM 2025 · 3 citations
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 2 citations
- MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in MemesSiddhant Agarwal, Adya Dhuler, Polly Ruhnke, Melvin Speisman et al.AAAI 2026
Builds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
Related papers
- Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in MemesXuanrui Lin, Chao Jia, Junhui Ji, Hui Han et al.WWW 2025 · 9 citations
- I know what you MEME! Understanding and Detecting Harmful Memes with Multimodal Large Language ModelsYong Zhuang, Keyan Guo, Juan Wang, Yiheng Jing et al.NDSS 2025
- Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal FusionYi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei et al.AAAI 2026
- MemeIntel: Explainable Detection of Propagandistic and Hateful MemesMohamed Bayan Kmainasi, Abul Hasnat, Md. Arid Hasan, Ali Ezzat Shahroor et al.EMNLP 2025 · 1 citation
- MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language ModelsZhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao et al.EMNLP 2025
