Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
Peng Qi, Zehong Yan, Wynne Hsu, Mong-Li Lee
Abstract
Misinformation is a prevalent societal issue due to its potential high risks. Out-Of-Context (OOC) misinformation, where authentic images are repurposed with false text, is one of the easiest and most effective ways to mislead audiences. Current methods focus on assessing image- text consistency but lack convincing explanations for their judgments, which are essential for debunking misinformation. While Multimodal Large Language Models (MLLMs) have rich knowledge and innate capability for visual rea- soning and explanation generation, they still lack sophisti- cation in understanding and discovering the subtle cross- modal differences. In this paper, we introduce Sniffer,a novel multimodal large language model specifically engi- neered for OOC misinformation detection and explanation. Snifferemploys two-stage instruction tuning on Instruct- BLIP. The first stage refines the model's concept alignment of generic objects with news-domain entities and the sec- ond stage leverages OOC-specific instruction data gener- ated by language-only GPT-4 to fine-tune the model's dis- criminatory powers. Enhanced by external tools and re- trieval, Sniffernot only detects inconsistencies between text and image but also utilizes external knowledge for con- textual verification. Our experiments show that Sniffersurpasses the original MLLM by over 40% and outperforms state-of-the-art methods in detection accuracy. Snifferalso provides accurate and persuasive explanations as val- idated by quantitative and human evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fca339b9-a485-4778-9a96-8a7ae78c0054Cited by top-tier papers26
- FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsXuannan Liu, Peipei Li, Huaibo Huang, Zekun Li et al.ACM MM 2024 · 46 citations
- Supernotes: Driving Consensus in Crowd-Sourced Fact-CheckingSoham De, Michiel A. Bakker, Jay Baxter, Martin SaveskiWWW 2025 · 31 citations
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human FeedbackJiaming Ji, Xinyu Chen, Rui Pan, Han Zhu et al.NeurIPS 2025 · 28 citations
- Fact-R1: Towards Explainable Video Misinformation Detection with Deep ReasoningFanrui Zhang, Dian Li, Qiang Zhang, Jun Chen et al.NeurIPS 2025 · 20 citations
- Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual SteganographySongze Li, Jiameng Cheng, Yiming Li, Xiaojun Jia et al.NDSS 2026 · 9 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal MediaGrace Luo, Trevor Darrell, Anna RohrbachEMNLP 2021 · 58 citations
- Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language ModelsRunhe Lai, Xinhua Lu, Yanqi Wu, Jinlun Ye et al.ICML 2026 · 1 citation
- Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual GroundingPinxue Guo, Chongruo Wu, Xinyu Zhou, Lingyi Hong et al.AAAI 2026
- Probabilistic Concept Graph Reasoning for Multimodal Misinformation DetectionRuichao Yang, Wei Gao, Xiaobin Zhu, Jing Ma et al.CVPR 2026 · 1 citation
- From Pixels to Semantics: A Novel MLLM-Driven Approach for Explainable Tampered Text DetectionGuitao Xu, Ziqi Yi, Peirong Zhang, Jiahuan Cao et al.ACM MM 2025 · 2 citations
