Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in Memes
Xuanrui Lin, Chao Jia, Junhui Ji, Hui Han, Usman Naseem
Abstract
Memes serve as a powerful medium of expression in the digital age, shaping cultural discourse and conveying ideas succinctly and engagingly. However, their potential for social abuse highlights the importance of developing effective methods to detect harmful content within memes. Recent studies on memes have focused on transforming images into textual captions using large language models (LLMs). However, these approaches often result in non-informative captions. Furthermore, previous methods have only been tested on limited datasets, providing insufficient evidence of their robustness. To address these limitations, we present a multimodal, agent-based framework designed to generate informative visual descriptions of memes by asking insightful questions to improve visual descriptions in zero-shot visual question-answering settings. Specifically, we leverage an LLM as agents with distinct roles and a large multimodal model (LMM) as a vision expert. These agents first analyze the images and then ask informative questions related to potential social abuse in memes to obtain high-quality answers about the images. Through continuous discussion guided by instructional prompts, the agents gather high-quality information while repeatedly acquiring image data from the LMM, which helps detect social abuse in memes. Finally, the discussion history and basic information are classified using the LLM to obtain the final prediction results in a zero-shot setting. Experimental results on a meme benchmark dataset sourced from 5 diverse meme datasets, comprising 6,626 memes spanning 5 tasks of varying complexity related to social abuse, demonstrate that our framework outperforms state-of-the-art methods, with detailed comparative analysis and ablation studies, further validating its generalizability and ability to retrieve more relevant information for detecting social abuse in memes. Disclaimer: This paper contains content that may be disturbing to some readers. CCS Concepts • Information Extraction → Multimodality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0c37725-99c8-4ab2-9efb-8c6267884e22Cited by top-tier papers3
- All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept ReproductionZiyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang et al.ACL 2026
- They Said Memes Were Harmless - We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural ReferencesSahil Tripathi, Gautam Siddharth Kashyap, Mehwish Nasim, Jian Yang et al.WWW 2026
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme UnderstandingWenqing Hou, Hongkui Tu, Ye Wang, Yue Zhang et al.ACL 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- Towards Low-Resource Harmful Meme Detection with LMM AgentsJianzhao Huang, Hongzhan Lin, Ziyan Liu, Ziyang Luo et al.EMNLP 2024 · 2 citations
- MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme InterventionPrince Jha, Raghav Jain, Konika Mandal, Aman Chadha et al.ACL 2024 · 5 citations
- AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on HarmfulnessZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo et al.ACL 2025
- Is Having Rationales Enough? Rethinking Knowledge Enhancement for Multimodal Hateful Meme DetectionJunyu Lu, Bo Xu, Xiaokun Zhang, Haohao Zhu et al.SIGIR 2025 · 3 citations
- MIND: A Multi-agent Framework for Zero-shot Harmful Meme DetectionZiyan Liu, Chunxiao Fan, Haoran Lou, Yuexin Wu et al.ACL 2025 · 15 citations
