Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
Bing Wang, Ximing Li, Yanjun Wang, Changchun Li, Lin Yuanbo Wu, Buyu Wang, Shengsheng Wang
Abstract
Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text modality may be much more informative than the image modality because the text generally describes the whole event/story of the current post but the image often presents partial scenes only. Our preliminary empirical results indicate that the image modality exactly contributes less to MMD. Upon this idea, we propose a new MMD method named RETSIMD. Specifically, we suppose that each text can be divided into several segments, and each text segment describes a partial scene that can be presented by an image. Accordingly, we split the text into a sequence of segments, and feed these segments into a pre-trained text-to-image generator to augment a sequence of images. We further incorporate two auxiliary objectives concerning text-image and image-label mutual information, and further post-train the generator over an auxiliary text-to-image generation benchmark dataset. Additionally, we propose a graph structure by defining three heuristic relationships between images, and use a graph neural network to generate the fused features. Extensive empirical results validate the effectiveness of RETSIMD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf05f8cb-112c-47ab-9a55-2e3736bdbc84Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 670 citations
Related papers
- Birds of a Feather: Enhancing Multimodal Fake News Detection Via Multi-Element RetrievalXueqin Chen, Xiaoyu Huang, Qiang Gao, Li Huang et al.ICDE 2025 · 2 citations
- Harmfully Manipulated Images Matter in Multimodal Misinformation DetectionBing Wang, Shengsheng Wang, Changchun Li, Renchu Guan et al.ACM MM 2024 · 5 citations
- Multimodal Taylor Series Network for Misinformation DetectionJiahao Sun, Chen Chen, Chunyan Hou, Yike Wu et al.WWW 2025 · 1 citation
- KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News DetectionPeican Zhu, Yubo Jing, Le Cheng, Keke Tang et al.ACM MM 2025 · 5 citations
- MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm DetectionHaochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou et al.CVPR 2026 · 4 citations
