VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
Yongkang Zhang, Dongyu She, Baiyu Ji, Qichuan Geng, Zhong Zhou, Yan Wang
摘要
The rapid evolution of generative AI, including such models as Sora, has intensified the threat of video misinformation. A critical challenge in detecting these AI-generated video misinformation lies in a fundamental disconnect between existing datasets and practical deception tactics. Current datasets often disrupt cross-modal consistency through editing techniques, resulting in unrealistic and easily detectable artifacts. By contrast, generative video misinformation strives for semantic consistency across modalities to remain realism. To address this gap, we introduce RAVM: the first Realistic AI-Generated Video Misinformation Detection Dataset. Unlike existing Video Misinformation Detection (VMD) datasets that are limited to single-source manipulations, RAVM encompasses multiple manipulation sources-Claim, Video, Audio, and Cross-Modal Manipulation-each incorporating diverse manipulation techniques to generate realistic AI-generated video misinformation. To achieve this, we introduce an agent-driven framework for generating realistic video misinformation. Furthermore, we propose an IEEG model that represents multimodal evidence, fact-checking results, and their dependencies as an evidence graph for interpretable detection of AI-generated video misinformation. Extensive experiments on RAVM reveal the vulnerability of existing Multimodal Large Language Models (MLLMs) in detecting AI-generated video misinformation, while the proposed IEEG achieves stateof-the-art performance on RAVM. The dataset is publicly
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang 等CVPR 2026 · 被引用 179 次
- FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspective of Creative ProcessYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi 等ACM MM 2024 · 被引用 45 次
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan 等ICLR 2026 · 被引用 38 次
相关 Paper
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual ContextsYuchen Zhang, Yaxiong Wang, Yujiao Wu, Lianwei Wu 等CVPR 2026 · 被引用 8 次
- Combating Online Misinformation Videos: Characterization, Detection, and Future DirectionsYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi 等ACM MM 2023 · 被引用 38 次
- From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation AnalysisZhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang 等ACL 2026
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsJinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu 等ACM MM 2025 · 被引用 2 次
- ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake DatasetKartik Thakral, Rishabh Ranjan, Akanksha Singh, Akshat Jain 等ICLR 2025
