Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in Memes
Xuanrui Lin, Chao Jia, Junhui Ji, Hui Han, Usman Naseem
摘要
Memes serve as a powerful medium of expression in the digital age, shaping cultural discourse and conveying ideas succinctly and engagingly. However, their potential for social abuse highlights the importance of developing effective methods to detect harmful content within memes. Recent studies on memes have focused on transforming images into textual captions using large language models (LLMs). However, these approaches often result in non-informative captions. Furthermore, previous methods have only been tested on limited datasets, providing insufficient evidence of their robustness. To address these limitations, we present a multimodal, agent-based framework designed to generate informative visual descriptions of memes by asking insightful questions to improve visual descriptions in zero-shot visual question-answering settings. Specifically, we leverage an LLM as agents with distinct roles and a large multimodal model (LMM) as a vision expert. These agents first analyze the images and then ask informative questions related to potential social abuse in memes to obtain high-quality answers about the images. Through continuous discussion guided by instructional prompts, the agents gather high-quality information while repeatedly acquiring image data from the LMM, which helps detect social abuse in memes. Finally, the discussion history and basic information are classified using the LLM to obtain the final prediction results in a zero-shot setting. Experimental results on a meme benchmark dataset sourced from 5 diverse meme datasets, comprising 6,626 memes spanning 5 tasks of varying complexity related to social abuse, demonstrate that our framework outperforms state-of-the-art methods, with detailed comparative analysis and ablation studies, further validating its generalizability and ability to retrieve more relevant information for detecting social abuse in memes. Disclaimer: This paper contains content that may be disturbing to some readers. CCS Concepts • Information Extraction → Multimodality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept ReproductionZiyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang 等ACL 2026
- They Said Memes Were Harmless - We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural ReferencesSahil Tripathi, Gautam Siddharth Kashyap, Mehwish Nasim, Jian Yang 等WWW 2026
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme UnderstandingWenqing Hou, Hongkui Tu, Ye Wang, Yue Zhang 等ACL 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
相关 Paper
- Towards Low-Resource Harmful Meme Detection with LMM AgentsJianzhao Huang, Hongzhan Lin, Ziyan Liu, Ziyang Luo 等EMNLP 2024 · 被引用 2 次
- MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme InterventionPrince Jha, Raghav Jain, Konika Mandal, Aman Chadha 等ACL 2024 · 被引用 5 次
- AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on HarmfulnessZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo 等ACL 2025
- Is Having Rationales Enough? Rethinking Knowledge Enhancement for Multimodal Hateful Meme DetectionJunyu Lu, Bo Xu, Xiaokun Zhang, Haohao Zhu 等SIGIR 2025 · 被引用 3 次
- MIND: A Multi-agent Framework for Zero-shot Harmful Meme DetectionZiyan Liu, Chunxiao Fan, Haoran Lou, Yuexin Wu 等ACL 2025 · 被引用 15 次
