Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models
Hongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma, Bo Wang, Ruichao Yang
摘要
The age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, which is not explicitly conveyed through the surface text and image. However, existing harmful meme detection methods do not present readable explanations that unveil such implicit meaning to support their detection decisions. In this paper, we propose an explainable approach to detect harmful memes, achieved through reasoning over conflicting rationales from both harmless and harmful positions. Specifically, inspired by the powerful capacity of Large Language Models (LLMs) on text generation and reasoning, we first elicit multimodal debate between LLMs to generate the explanations derived from the contradictory arguments. Then we propose to fine-tune a small language model as the debate judge for harmfulness inference, to facilitate multimodal fusion between the harmfulness rationales and the intrinsic multimodal information within memes. In this way, our model is empowered to perform dialectical reasoning over intricate and implicit harm-indicative patterns, utilizing multimodal explanations originating from both harmless and harmful arguments. Extensive experiments on three public meme datasets demonstrate that our harmful meme detection approach achieves much better performance than state-of-the-art methods and exhibits a superior capacity for explaining the meme harmfulness of the model predictions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- MIND: A Multi-agent Framework for Zero-shot Harmful Meme DetectionZiyan Liu, Chunxiao Fan, Haoran Lou, Yuexin Wu 等ACL 2025 · 被引用 15 次
- Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate DetectionJian Lang, Rongpei Hong, Jin Xu, Yili Li 等WWW 2025 · 被引用 14 次
- Multi-Granular Multimodal Clue Fusion for Meme UnderstandingLi Zheng, Hao Fei, Ting Dai, Zuquan Peng 等AAAI 2025 · 被引用 13 次
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng 等ACL 2024 · 被引用 10 次
- ExPO-HM: Learning to Explain-then-Detect for Hateful Meme DetectionJingbiao Mei, Mingsheng Sun, Jinghong Chen, Pengda Qin 等ICLR 2026 · 被引用 8 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Towards Low-Resource Harmful Meme Detection with LMM AgentsJianzhao Huang, Hongzhan Lin, Ziyan Liu, Ziyang Luo 等EMNLP 2024 · 被引用 2 次
- Is Having Rationales Enough? Rethinking Knowledge Enhancement for Multimodal Hateful Meme DetectionJunyu Lu, Bo Xu, Xiaokun Zhang, Haohao Zhu 等SIGIR 2025 · 被引用 3 次
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 被引用 2 次
- AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on HarmfulnessZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo 等ACL 2025
- Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal FusionYi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei 等AAAI 2026
