Cross-modal Contrastive Learning for Multimodal Fake News Detection
Longzheng Wang, Chuang Zhang, Hongbo Xu, Yongxiu Xu, Xiaohan Xu, Siqi Wang
Abstract
Automatic detection of multimodal fake news has gained a widespread attention recently. Many existing approaches seek to fuse unimodal features to produce multimodal news representations. However, the potential of powerful cross-modal contrastive learning methods for fake news detection has not been well exploited. Besides, how to aggregate features from different modalities to boost the performance of the decision-making process is still an open question. To address that, we propose COOLANT, a crossmodal contrastive learning framework for multimodal fake news detection, aiming to achieve more accurate image-text alignment. To further capture the fine-grained alignment between vision and language, we leverage an auxiliary task to soften the loss term of negative samples during the contrast process. A cross-modal fusion module is developed to learn the cross-modality correlations. An attention mechanism with an attention guidance module is implemented to help effectively and interpretably aggregate the aligned unimodal representations and the cross-modality correlations. Finally, we evaluate the COOLANT and conduct a comparative study on two widely used datasets, Twitter and Weibo. The experimental results demonstrate that our COOLANT outperforms previous approaches by a large margin and achieves new state-of-the-art results on the two datasets. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c742b17c-52b7-4466-ade5-b82494363236Cited by top-tier papers10
- Each Fake News Is Fake in Its Own Way: An Attribution Multi-Granularity Benchmark for Multimodal Fake News DetectionHao Guo, Zihan Ma, Zhi Zeng, Minnan Luo et al.AAAI 2025 · 7 citations
- Harmfully Manipulated Images Matter in Multimodal Misinformation DetectionBing Wang, Shengsheng Wang, Changchun Li, Renchu Guan et al.ACM MM 2024 · 5 citations
- FactGuard: Agentic Video Misinformation Detection via Reinforcement LearningZehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng et al.ICML 2026 · 3 citations
- Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation DetectorsBing Wang, Ximing Li, Mengzhe Ye, Changchun Li et al.ACM MM 2025 · 2 citations
- Probabilistic Concept Graph Reasoning for Multimodal Misinformation DetectionRuichao Yang, Wei Gao, Xiaobin Zhu, Jing Ma et al.CVPR 2026 · 1 citation
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- External Reliable Information-enhanced Multimodal Contrastive Learning for Fake News DetectionBiwei Cao, Qihang Wu, Jiuxin Cao, Bo Liu et al.AAAI 2025 · 11 citations
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui et al.WWW 2022 · 325 citations
- Entity Graph Alignment and Visual Reasoning for Multimodal Fake News DetectionGuoyi Li, Die Hu, Xiaomeng Fu, Qirui Tang et al.ACM MM 2025 · 2 citations
- Knowledge-Enhanced Multimodal Fake News Detection: Semantic Visual and Priority FusionQin Zhang, Jiaying Liu, Qian Tao, Zhiwei Guo et al.WWW 2026
- Retrieval-Augmented Multimodal Model for Fake News DetectionYiheng Li, Weihai Lu, Hanyi Yu, Yue WangSIGIR 2026 · 4 citations
