Invariant Meets Specific: A Scalable Harmful Memes Detection Framework
Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu
Abstract
Harmful memes detection is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. Current research on this task mainly focuses on multimodal dual-stream models. However, the existing works ignore the misalignment of the memes caused by the modality gap. Moreover, the cross-modal interaction in the dual-stream models is insufficient to identify harmful memes. To this end, this paper proposes a scalable invariant and specific modality (ISM) representations framework via graph neural networks. The proposed ISM framework provides a comprehensive and disentangled view for memes and promotes inter-modal interaction. Specifically, ISM projects each modality to two distinct spaces. The first space is modality-invariant, learning the corresponding commonalities and reducing the modality gap. The second space is modality-specific, holding the distinctive characteristics of each modality and complementing the common latent features captured in invariant spaces. Then, we construct fully connected visual and textual graphs for each space. The unimodal graphs are fused to dynamically balance inter-modal and intra-modal relationships, which are complementary to the dual-stream models. Finally, an adaptive module is designed to weigh the proportion of each fusion graph for memes. Moreover, the mainstream multimodal dual-stream models could be employed as the backbone flexibly. Extensive experiments on five publicly available datasets show that the proposed ISM provides a stable improvement over baselines and produces a competitive performance compared with the existing harmful memes detection methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3e472986-d34e-4b49-ba30-0f94277dbdbfCited by top-tier papers4
- Multi-Granular Multimodal Clue Fusion for Meme UnderstandingLi Zheng, Hao Fei, Ting Dai, Zuquan Peng et al.AAAI 2025 · 13 citations
- They Said Memes Were Harmless - We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural ReferencesSahil Tripathi, Gautam Siddharth Kashyap, Mehwish Nasim, Jian Yang et al.WWW 2026
- Uncertainty-Guided Modal Rebalance for Hateful Memes DetectionChuanpeng Yang, Yaxin Liu, Fuqing Zhu, Jizhong Han et al.ACL 2024
- From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-ImprovementJian Lang, Rongpei Hong, Ting Zhong, Leiting Chen et al.KDD 2026
Related papers
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme UnderstandingWenqing Hou, Hongkui Tu, Ye Wang, Yue Zhang et al.ACL 2026
- Disentangling Hate in Online MemesRoy Ka-Wei Lee, Rui Cao, Ziqing Fan, Jing Jiang et al.ACM MM 2021 · 85 citations
- DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm DetectionZhihong Zhu, Kefan Shen, Zhaorun Chen, Yunyan Zhang et al.EMNLP 2024 · 5 citations
- TOT:Topology-Aware Optimal Transport for Multimodal Hate DetectionLinhao Zhang, Li Jin, Xian Sun, Guangluan Xu et al.AAAI 2023 · 9 citations
- Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionSijie Mai, Haifeng Hu, Songlong XingAAAI 2020 · 233 citations
