Your tone speaks louder than your face! Modality Order Infused Multi-modal Sarcasm Detection
Mohit Tomar, Abhisek Tiwari, Tulika Saha, Sriparna Saha
摘要
Figurative language is an essential component of human communication, and detecting sarcasm in text has become a challenging yet highly popular task in natural language processing. As humans, we rely on a combination of visual and auditory cues, such as facial expressions and tone of voice, to comprehend a message. Our brains are implicitly trained to integrate information from multiple senses to form a complete understanding of the message being conveyed, a process known as multi-sensory integration. The combination of different modalities not only provides additional information but also amplifies the information conveyed by each modality in relation to the others. Thus, the infusion order of different modalities also plays a significant role in multimodal processing. In this paper, we investigate the impact of different modality infusion orders for identifying sarcasm in dialogues. We propose a modality order-driven module integrated into a transformer network, MO-Sarcation that fuses modalities in an ordered manner. Our model outperforms several state-of-the-art models by 1-3% across various metrics, demonstrating the crucial role of modality order in sarcasm detection. The obtained improvements and detailed analysis show that audio tone should be infused with textual content, followed by visual information to identify sarcasm efficiently. The code and dataset are available at https://github.com/mohit2b/MO-Sarcation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal FusionSijie Mai, Shiqin HanCVPR 2026 · 被引用 1 次
- SAGE: Synergistic Adaptive Gating of Experts for Hateful Video DetectionJie Huang, Xin Liao, Junjie Wang, Mingyang Li 等ACL 2026
- MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in DialoguesDiandian Guo, Fangfang Yuan, Cong Cao, Xixun Lin 等WWW 2026
- Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm DetectionYihua Wang, Qi Jia, Cong Xu, Feiyu Chen 等AAAI 2026
- Supervised Attention Mechanism for Low-quality Multimodal DataSijie Mai, Shiqin Han, Haifeng HuEMNLP 2025
相关 Paper
- When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party DialoguesShivani Kumar, Atharva Kulkarni, Md. Shad Akhtar, Tanmoy ChakrabortyACL 2022 · 被引用 54 次
- Explaining (Sarcastic) Utterances to Enhance Affect Understanding in Multimodal DialoguesShivani Kumar, Ishani Mondal, Md. Shad Akhtar, Tanmoy ChakrabortyAAAI 2023 · 被引用 22 次
- Multimodal Sarcasm Target Identification in TweetsJiquan Wang, Lin Sun, Yi Liu, Meizhi Shao 等ACL 2022 · 被引用 28 次
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 被引用 91 次
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 被引用 49 次
