From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation Analysis
Zhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang, Yifei Yang, Xiangzheng Kong, Herun Wan, Zihan Ma
摘要
Video misinformation detection is often approached as a binary veracity classification problem, overlooking the complex reasoning required to explain how and why content misleads. Existing benchmarks fail to capture the diversity of manipulation strategies, such as AI-generated edits and out-of-context manipulation, and do not evaluate whether models can provide process-level justifications for their judgments. We address these limitations with MISVIDEOQA, a multi-turn benchmark designed to assess comprehensive understanding and reasoning in video misinformation analysis. MISVIDEOQA covers 12 fine-grained video categories and evaluates models along six dimensions, progressing from perceptual attribution to intent and persuasion analysis. Recognizing that standard MLLMs struggle to sustain such structured, evidence-based deduction, we propose MISAGENT, a Delphiinspired multi-agent framework in which specialized agents collaboratively integrate multimodal cues with external evidence. Experimental results show that state-of-the-art multimodal large language models perform poorly on MISVIDEOQA, while MISAGENT consistently improves reasoning accuracy and explanation quality. Together, our benchmark and framework establish a unified foundation for reliable, interpretable, and evidence-grounded video misinformation analysis. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang 等ICML 2024 · 被引用 182 次
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal MediaGrace Luo, Trevor Darrell, Anna RohrbachEMNLP 2021 · 被引用 58 次
- Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation DetectionPeng Qi, Zehong Yan, Wynne Hsu, Mong-Li LeeCVPR 2024 · 被引用 54 次
- FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspective of Creative ProcessYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi 等ACM MM 2024 · 被引用 45 次
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu 等ACM MM 2024 · 被引用 43 次
相关 Paper
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang 等WWW 2026 · 被引用 3 次
- FactGuard: Agentic Video Misinformation Detection via Reinforcement LearningZehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng 等ICML 2026 · 被引用 3 次
- AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMsShuhan Xia, Peipei Li, Xuannan Liu, Dongsen Zhang 等CVPR 2026 · 被引用 1 次
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li 等ICLR 2026 · 被引用 22 次
- Perception, Understanding and Reasoning: A Multimodal Benchmark for Video Fake News DetectionYakun Cui, Peng Qi, Fushuo Huo, Hang Du 等ACL 2026
