Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
Fanrui Zhang, Dian Li, Qiang Zhang, Jun Chen, Sinbadliu, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha
摘要
The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive content. To address these challenges, we introduce FakeVV, a large-scale benchmark comprising over 100,000 video-text pairs with fine-grained, interpretable annotations. In addition, we further propose Fact-R1, a novel framework that integrates deep reasoning with collaborative rule-based reinforcement learning. Fact-R1 is trained through a three-stage process: (1) misinformation long-Chain-of-Thought (CoT) instruction tuning, (2) preference alignment via Direct Preference Optimization (DPO), and (3) Group Relative Policy Optimization (GRPO) using a novel verifiable reward function. This enables Fact-R1 to exhibit emergent reasoning behaviors comparable to those observed in advanced text-based reinforcement learning systems, but in the more complex multimodal misinformation setting. Our work establishes a new paradigm for misinformation detection, bridging large-scale video understanding, reasoning-guided alignment, and interpretable verification. * Equal contribution. Work done during internship at Tencent QQ, as a part of QQ MLLM project. † Corresponding author. ‡ Project leader of QQ MLLM project. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang 等WWW 2026 · 被引用 3 次
- FactGuard: Agentic Video Misinformation Detection via Reinforcement LearningZehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng 等ICML 2026 · 被引用 3 次
- Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level GuidanceQiang Zhang, Fanrui Zhang, Jiawei Liu, Ming Hu 等NeurIPS 2025 · 被引用 1 次
- From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation AnalysisZhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang 等ACL 2026
- VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video MisinformationYongkang Zhang, Dongyu She, Baiyu Ji, Qichuan Geng 等CVPR 2026
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui 等WWW 2022 · 被引用 325 次
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal MediaGrace Luo, Trevor Darrell, Anna RohrbachEMNLP 2021 · 被引用 58 次
- Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation DetectionPeng Qi, Zehong Yan, Wynne Hsu, Mong-Li LeeCVPR 2024 · 被引用 54 次
相关 Paper
- Perception, Understanding and Reasoning: A Multimodal Benchmark for Video Fake News DetectionYakun Cui, Peng Qi, Fushuo Huo, Hang Du 等ACL 2026
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake DetectionTuan Nguyen, Naseem Khan, Khang Tran, Hai Phan 等ICML 2026
- Entity Graph Alignment and Visual Reasoning for Multimodal Fake News DetectionGuoyi Li, Die Hu, Xiaomeng Fu, Qirui Tang 等ACM MM 2025 · 被引用 2 次
- VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningHao Tan, jun lan, Senyuan Shi, Zichang Tan 等ICML 2026 · 被引用 12 次
