Multimodal Sarcasm Target Identification in Tweets
Jiquan Wang, Lin Sun, Yi Liu, Meizhi Shao, Zengwei Zheng
Abstract
Sarcasm is important to sentiment analysis on social media. Sarcasm Target Identification (STI) deserves further study to understand sarcasm in depth. However, text lacking context or missing sarcasm target makes target identification very difficult. In this paper, we introduce multimodality to STI and present Multimodal Sarcasm Target Identification (MSTI) task. We propose a novel multi-scale cross-modality model that can simultaneously perform textual target labeling and visual target detection. In the model, we extract multi-scale visual features to enrich spatial information for different sized visual sarcasm targets. We design a set of convolution networks to unify multi-scale visual features with textual features for cross-modal attention learning, and correspondingly a set of transposed convolution networks to restore multi-scale visual information. The results show that visual clues can improve the performance of TSTI by a large margin, and VSTI achieves good accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- FigMemes: A Dataset for Figurative Language Identification in Politically-Opinionated MemesChen Liu, Gregor Geigle, Robin Krebs, Iryna GurevychEMNLP 2022 · 15 citations
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng et al.ACL 2024 · 10 citations
- DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm UnderstandingHang Du, Guoshun Nan, Sicheng Zhang, Binzhu Xie et al.AAAI 2024 · 9 citations
- Well, Now We Know! Unveiling Sarcasm: Initiating and Exploring Multimodal Conversations with ReasoningGopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal et al.AAAI 2024 · 6 citations
- DIP: Dual Incongruity Perceiving Network for Sarcasm DetectionChangsong Wen, Guoli Jia, Jufeng YangCVPR 2023
Builds on3
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 153 citations
Related papers
- MSTI-Plus: Introducing Non-Sarcasm Reference Materials to Enhance Multimodal Sarcasm Target IdentificationFengmao Lv, Mengting Xiong, Junlin Fang, Lingli Zhang et al.WWW 2025 · 1 citation
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui et al.ACM MM 2021 · 128 citations
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen et al.AAAI 2023 · 84 citations
- Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional NetworkBin Liang, Chenwei Lou, Xiang Li, Min Yang et al.ACL 2022 · 151 citations
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 91 citations
