From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation Analysis
Zhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang, Yifei Yang, Xiangzheng Kong, Herun Wan, Zihan Ma
Abstract
Video misinformation detection is often approached as a binary veracity classification problem, overlooking the complex reasoning required to explain how and why content misleads. Existing benchmarks fail to capture the diversity of manipulation strategies, such as AI-generated edits and out-of-context manipulation, and do not evaluate whether models can provide process-level justifications for their judgments. We address these limitations with MISVIDEOQA, a multi-turn benchmark designed to assess comprehensive understanding and reasoning in video misinformation analysis. MISVIDEOQA covers 12 fine-grained video categories and evaluates models along six dimensions, progressing from perceptual attribution to intent and persuasion analysis. Recognizing that standard MLLMs struggle to sustain such structured, evidence-based deduction, we propose MISAGENT, a Delphiinspired multi-agent framework in which specialized agents collaboratively integrate multimodal cues with external evidence. Experimental results show that state-of-the-art multimodal large language models perform poorly on MISVIDEOQA, while MISAGENT consistently improves reasoning accuracy and explanation quality. Together, our benchmark and framework establish a unified foundation for reliable, interpretable, and evidence-grounded video misinformation analysis. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41f77e29-8175-4150-92ff-f3ce64b425bbBuilds on21
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang et al.ICML 2024 · 182 citations
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal MediaGrace Luo, Trevor Darrell, Anna RohrbachEMNLP 2021 · 58 citations
- Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation DetectionPeng Qi, Zehong Yan, Wynne Hsu, Mong-Li LeeCVPR 2024 · 54 citations
- FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspective of Creative ProcessYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi et al.ACM MM 2024 · 45 citations
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu et al.ACM MM 2024 · 43 citations
Related papers
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang et al.WWW 2026 · 3 citations
- FactGuard: Agentic Video Misinformation Detection via Reinforcement LearningZehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng et al.ICML 2026 · 3 citations
- AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMsShuhan Xia, Peipei Li, Xuannan Liu, Dongsen Zhang et al.CVPR 2026 · 1 citation
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li et al.ICLR 2026 · 22 citations
- Perception, Understanding and Reasoning: A Multimodal Benchmark for Video Fake News DetectionYakun Cui, Peng Qi, Fushuo Huo, Hang Du et al.ACL 2026
