Dynamic Routing Transformer Network for Multimodal Sarcasm Detection
Yuan Tian, Nan Xu, Ruike Zhang, Wenji Mao
摘要
Multimodal sarcasm detection is an important research topic in natural language processing and multimedia computing, and benefits a wide range of applications in multiple domains. Most existing studies regard the incongruity between image and text as the indicative clue in identifying multimodal sarcasm. To capture cross-modal incongruity, previous methods rely on fixed architectures in network design, which restricts the model from dynamically adjusting to diverse image-text pairs. Inspired by routing-based dynamic network, we model the dynamic mechanism in multimodal sarcasm detection and propose the Dynamic Routing Transformer Network (DynRT-Net). Our method utilizes dynamic paths to activate different routing transformer modules with hierarchical co-attention adapting to cross-modal incongruity. Experimental results on a public dataset demonstrate the effectiveness of our method compared to the state-of-the-art methods. Our codes are available at https://github.com/TIAN-viola/DynRT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction ExpertsHaofei Yu, Zhengyang Qi, Lawrence Jang, Russ Salakhutdinov 等EMNLP 2024 · 被引用 11 次
- MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm DetectionYifeng Xie, Zhihong Zhu, Xin Chen, Zhanpeng Chen 等ACM MM 2024 · 被引用 11 次
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng 等ACL 2024 · 被引用 10 次
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu 等EMNLP 2024 · 被引用 10 次
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai 等ACL 2025 · 被引用 5 次
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao 等SIGIR 2021 · 被引用 187 次
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 被引用 153 次
- Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional NetworkBin Liang, Chenwei Lou, Xiang Li, Min Yang 等ACL 2022 · 被引用 151 次
- TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringYiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun 等ICCV 2021 · 被引用 128 次
相关 Paper
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen 等AAAI 2023 · 被引用 84 次
- DIP: Dual Incongruity Perceiving Network for Sarcasm DetectionChangsong Wen, Guoli Jia, Jufeng YangCVPR 2023
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 被引用 91 次
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui 等ACM MM 2021 · 被引用 128 次
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 被引用 49 次
