DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm Detection
Zhihong Zhu, Kefan Shen, Zhaorun Chen, Yunyan Zhang, Yuyan Chen, Xiaoqi Jiao, Zhongwei Wan, Shaorong Xie, Wei Liu, Xian Wu, Yefeng Zheng
Abstract
Capturing inter-modal incongruities within the text-image pair is a critical challenge in multimodal sarcasm detection (MSD). Fortunately, graph neural networks (GNNs) have made promising advancements in MSD, which show advantages in explicitly capturing data relationships. Nevertheless, current GNN-based MSD methods do not effectively address some of the inherent limitations of GNNs, which include: 1) neglecting high-order relationships, and 2) underestimating high-frequency messages. In this paper, we propose a Dual Graph-based Learning Framework (DGLF) to address the above two issues. Specifically, we construct a hypergraph to perform high-order aware propagation and a vanilla graph to perform highfrequency enhanced propagation, respectively. We empower GNNs to 1) better capture the inherent and complicated relationships based on the hypergraph and 2) deliver sufficient modeling through high-frequency enhanced messages on the vanilla graph. Besides, we introduce multi-modal fusion information bottleneck to effectively fuse the two learned graph features. Experimental results on two benchmark datasets show that the proposed model outperforms previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 827db003-1b05-4555-b5fe-b5a31671fa8fBuilds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Beyond Low-frequency Information in Graph Convolutional NetworksDeyu Bo, Xiao Wang, Chuan Shi, Huawei ShenAAAI 2021 · 773 citations
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 153 citations
Related papers
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 91 citations
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen et al.AAAI 2023 · 84 citations
- Invariant Meets Specific: A Scalable Harmful Memes Detection FrameworkChuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin HuACM MM 2023 · 9 citations
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui et al.ACM MM 2021 · 128 citations
- DIP: Dual Incongruity Perceiving Network for Sarcasm DetectionChangsong Wen, Guoli Jia, Jufeng YangCVPR 2023
