Multimodal Taylor Series Network for Misinformation Detection
Jiahao Sun, Chen Chen, Chunyan Hou, Yike Wu, Xiaojie Yuan
Abstract
With the rapid development of the Internet and the widespread use of social media, the proliferation of multimodal misinformation combining images and text poses serious risks to societal trust, individual well-being, and the integrity of AI models trained on such data. Recently, the automatic detection multimodal misinformation has become an essential area of research. However, traditional methods often rely on hierarchical neural networks that compress and fuse modalities, potentially overlooking deeper interactions between modalities and reducing model interpretability. In this paper, we present a novel Multimodal Taylor Series (MTS) network for detecting multimodal misinformation. The MTS network leverages Taylor series expansion to explicitly capture both low-order and high-order interactions between modalities, which also enhances interpretability by decomposing the model's processing into distinct terms. Additionally, the proposed MTS network avoids exponential parameter growth and maintains linear scalability, allowing the model to effectively capture complex cross-modal correlations. Extensive experiments on three benchmark datasets demonstrate that the MTS network significantly outperforms state-of-the-art models. We will release our code after the final publication of the paper. CCS Concepts • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74ea90f2-1f4b-493d-abd6-ccfd87654f84Cited by top-tier papers1
Ask how each one uses itBuilds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- Mining Dual Emotion for Fake News DetectionXueyao Zhang, Juan Cao, Xirong Li, Qiang Sheng et al.WWW 2021 · 332 citations
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui et al.WWW 2022 · 325 citations
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 153 citations
Related papers
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality PerspectiveBing Wang, Ximing Li, Yanjun Wang, Changchun Li et al.AAAI 2026 · 1 citation
- TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation DetectionZehong Yan, Peng Qi, Wynne Hsu, Mong-Li LeeEMNLP 2025 · 1 citation
- KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News DetectionPeican Zhu, Yubo Jing, Le Cheng, Keke Tang et al.ACM MM 2025 · 5 citations
- Hierarchical Multi-modal Contextual Attention Network for Fake News DetectionShengsheng Qian, Jinguang Wang, Jun Hu, Quan Fang et al.SIGIR 2021 · 273 citations
- SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal ModelZhenglin Huang, Jinwei Hu, Xiangtai Li, Yiwei He et al.CVPR 2025
