Dynamic Routing Transformer Network for Multimodal Sarcasm Detection
Yuan Tian, Nan Xu, Ruike Zhang, Wenji Mao
Abstract
Multimodal sarcasm detection is an important research topic in natural language processing and multimedia computing, and benefits a wide range of applications in multiple domains. Most existing studies regard the incongruity between image and text as the indicative clue in identifying multimodal sarcasm. To capture cross-modal incongruity, previous methods rely on fixed architectures in network design, which restricts the model from dynamically adjusting to diverse image-text pairs. Inspired by routing-based dynamic network, we model the dynamic mechanism in multimodal sarcasm detection and propose the Dynamic Routing Transformer Network (DynRT-Net). Our method utilizes dynamic paths to activate different routing transformer modules with hierarchical co-attention adapting to cross-modal incongruity. Experimental results on a public dataset demonstrate the effectiveness of our method compared to the state-of-the-art methods. Our codes are available at https://github.com/TIAN-viola/DynRT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09f4ee72-f1db-4c29-ab1a-d1d4bfb58261Cited by top-tier papers7
- MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction ExpertsHaofei Yu, Zhengyang Qi, Lawrence Jang, Russ Salakhutdinov et al.EMNLP 2024 · 11 citations
- MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm DetectionYifeng Xie, Zhihong Zhu, Xin Chen, Zhanpeng Chen et al.ACM MM 2024 · 11 citations
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng et al.ACL 2024 · 10 citations
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu et al.EMNLP 2024 · 10 citations
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai et al.ACL 2025 · 5 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 153 citations
- Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional NetworkBin Liang, Chenwei Lou, Xiang Li, Min Yang et al.ACL 2022 · 151 citations
- TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringYiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun et al.ICCV 2021 · 128 citations
Related papers
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen et al.AAAI 2023 · 84 citations
- DIP: Dual Incongruity Perceiving Network for Sarcasm DetectionChangsong Wen, Guoli Jia, Jufeng YangCVPR 2023
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 91 citations
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui et al.ACM MM 2021 · 128 citations
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 49 citations
