M3D: MultiModal MultiDocument Fine-Grained Inconsistency Detection
Chia-Wei Tang, Ting-Chih Chen, Kiet Nguyen, Kazi Sajeed Mehrab, Alvi Md. Ishmam, Chris Thomas
Abstract
Validating claims from misinformation is a highly challenging task that involves understanding how each factual assertion within the claim relates to a set of trusted source materials. Existing approaches often make coarse-grained predictions but fail to identify the specific aspects of the claim that are troublesome and the specific evidence relied upon. In this paper, we introduce a method and new benchmark for this challenging task. Our method predicts the fine-grained logical relationship of each aspect of the claim from a set of multimodal documents, which include text, image(s), video(s), and audio(s). We also introduce a new benchmark (M^3DC) of claims requiring multimodal multidocument reasoning, which we construct using a novel claim synthesis technique. Experiments show that our approach significantly outperforms state-of-the-art baselines on this challenging task on two benchmarks while providing finer-grained predictions, explanations, and evidence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 345beeed-e1d8-4ebb-9de6-c9b31a2baff0Cited by top-tier papers2
- VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-CheckingMark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna RohrbachACL 2026 · 4 citations
- DEFAME: Dynamic Evidence-based FAct-checking with Multimodal ExpertsTobias Braun, Mark Rothermel, Marcus Rohrbach, Anna RohrbachICML 2025
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- Multimodal Fact-Level Attribution for Verifiable ReasoningDavid Wan, Han Wang, Ziyang Wang, Elias Stengel-Eskin et al.ICML 2026 · 2 citations
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu et al.ICLR 2026 · 53 citations
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in VideoHanoona Abdul Rasheed, Abdelrahman M Shaker, Anqi Tang, Muhammad Maaz et al.ICLR 2026 · 16 citations
- MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm DetectionHaochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou et al.CVPR 2026 · 4 citations
- MetaLogic: Logical Reasoning Explanations with Fine-Grained StructureYinya Huang, Hongming Zhang, Ruixin Hong, Xiaodan Liang et al.EMNLP 2022 · 3 citations
