Open-Domain, Content-based, Multi-modal Fact-checking of Out-of-Context Images via Online Resources
Sahar Abdelnabi, Rakibul Hasan, Mario Fritz
Abstract
Misinformation is now a major problem due to its poten-tial high risks to our core democratic and societal values and orders. Out-of-context misinformation is one of the easiest and effective ways used by adversaries to spread vi-ral false stories. In this threat, a real image is re-purposed to support other narratives by misrepresenting its context and/or elements. The internet is being used as the go-to way to verify information using different sources and modali-ties. Our goal is an inspectable method that automates this time-consuming and reasoning-intensive process by fact-checking the image-caption pairing using Web evidence. To integrate evidence and cues from both modalities, we intro-duce the concept of ‘multi-modal cycle-consistency check’ starting from the image/caption, we gather tex-tual/visual evidence, which will be compared against the other paired caption/image, respectively. Moreover, we propose a novel architecture, Consistency-Checking Network (CCN), that mimics the layered human reasoning across the same and different modalities: the caption vs. textual evidence, the image vs. visual evidence, and the image vs. caption. Our work offers the first step and bench-mark for open-domain, content-based, multi-modal fact-checking, and significantly outperforms previous baselines that did not leverage external evidence <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> For code, checkpoints, and dataset, check: https://s-abdelnabi.github.io/OoC-multi-modal-fc/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Bootstrapping Multi-View Representations for Fake News DetectionQichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian et al.AAAI 2023 · 111 citations
- Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation DetectionPeng Qi, Zehong Yan, Wynne Hsu, Mong-Li LeeCVPR 2024 · 54 citations
- Combating Online Misinformation Videos: Characterization, Detection, and Future DirectionsYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi et al.ACM MM 2023 · 38 citations
- Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for MisinformationMax Glockner, Yufang Hou, Iryna GurevychEMNLP 2022 · 23 citations
- ESCNet: Entity-enhanced and Stance Checking Network for Multi-modal Fact-CheckingFanrui Zhang, Jiawei Liu, Jingyi Xie, Qiang Zhang et al.WWW 2024 · 18 citations
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal MediaGrace Luo, Trevor Darrell, Anna RohrbachEMNLP 2021 · 58 citations
- Detecting Cross-Modal Inconsistency to Defend Against Neural Fake NewsReuben Tan, Bryan A. Plummer, Kate SaenkoEMNLP 2020 · 9 citations
- DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia et al.ACL 2020 · 7 citations
- Where Are the Facts? Searching for Fact-checked Information to Alleviate the Spread of Fake NewsNguyen Vo, Kyumin LeeEMNLP 2020 · 4 citations
Related papers
- Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and DisinformationXin Yuan, Jie Guo, Weidong Qiu, Zheng Huang et al.EMNLP 2023 · 11 citations
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park et al.ICCV 2021 · 56 citations
- ECENet: Explainable and Context-Enhanced Network for Muti-modal Fact verificationFanrui Zhang, Jiawei Liu, Qiang Zhang, Esther Sun et al.ACM MM 2023 · 23 citations
- "Image, Tell me your story!" Predicting the original meta-context of visual misinformationJonathan Tonglet, Marie-Francine Moens, Iryna GurevychEMNLP 2024 · 6 citations
- DEFAME: Dynamic Evidence-based FAct-checking with Multimodal ExpertsTobias Braun, Mark Rothermel, Marcus Rohrbach, Anna RohrbachICML 2025
