Retrieval-Augmented Multimodal Model for Fake News Detection
Yiheng Li, Weihai Lu, Hanyi Yu, Yue Wang
Abstract
In recent years, multi-modal multi-domain fake news detection has garnered increasing attention. Nevertheless, this direction presents two significant challenges: (1) Failure to Capture Cross-Instance Narrative Consistency: existing models usually evaluate each news in isolation, fail to capture cross-instance narrative consistency, and thus struggle to address the spread of cluster-based fake news driven by social media; (2) Lack of Domain-Specific Knowledge for Reasoning: conventional models, which rely solely on knowledge encoded in their parameters during training, struggle to generalize to new or data-scarce domains (e.g., emerging events or niche topics). To tackle these challenges, we introduce Retrieval-Augmented Multimodal Model for Fake News Detection (RAMM). First, RAMM employs a Multimodal Large Language Model (MLLM) as its backbone to capture cross-modal semantic information from news samples. Second, RAMM incorporates an Abstract Narrative Alignment Module. This component adaptively extracts abstract narrative consistency from diverse instances across distinct domains, aggregates relevant knowledge, and thereby enables the modeling of high-level narrative information. Finally, RAMM introduces a Semantic Representation Alignment Module, which aligns the model's decision-making paradigm with that of humans—specifically, it shifts the model's reasoning process from direct inference on multimodal features to an instance-based analogical reasoning process. Extensive experimental results on three public datasets validate the efficacy of our proposed approach. Our code is available at the following link: https://github.com/li-yiheng/RAMM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4ad3ccc-394a-4c5c-902f-2c43025cae36Builds on25
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui et al.WWW 2022 · 325 citations
- Embracing Domain Differences in Fake News: Cross-domain Fake News Detection using Multi-modal DataAmila Silva, Ling Luo, Shanika Karunasekera, Christopher LeckieAAAI 2021 · 170 citations
- Explainable Fake News Detection with Large Language Model via Defense Among Competing WisdomBo Wang, Jing Ma, Hongzhan Lin, Zhiwei Yang et al.WWW 2024 · 104 citations
Related papers
- Toward Multimodal Fake News Detection by Multi-perspective Rationale Generation and VerificationJunyang Chen, Yueqian Li, Ka Chung Ng, Huan Wang et al.AAAI 2026
- Entity Graph Alignment and Visual Reasoning for Multimodal Fake News DetectionGuoyi Li, Die Hu, Xiaomeng Fu, Qirui Tang et al.ACM MM 2025 · 2 citations
- See How You Read? Multi-Reading Habits Fusion Reasoning for Multi-Modal Fake News DetectionLianwei Wu, Pusheng Liu, Yanning ZhangAAAI 2023 · 40 citations
- DAMMFND: Domain-Aware Multimodal Multi-view Fake News DetectionWeihai Lu, Yu Tong, Zhiqiu YeAAAI 2025 · 21 citations
- MMDFND: Multi-modal Multi-Domain Fake News DetectionYu Tong, Weihai Lu, Zhe Zhao, Song Lai et al.ACM MM 2024 · 42 citations
