On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
Xiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li, Lei Bai, Xiaoming Liu, Guangtao Zhai, Xiaohong Liu
Abstract
Large numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identify diffusion-generated content with a diverse range of semantics. To advance the field of video forensics, we propose an innovative algorithm named Multi-Modal Detection(MM-Det) for detecting diffusion-generated videos. MM-Det utilizes the profound perceptual and comprehensive abilities of Large Multi-modal Models (LMMs) by generating a Multi-Modal Forgery Representation (MMFR) from LMM's multi-modal space, enhancing its ability to detect unseen forgery content. Besides, MM-Det leverages an In-and-Across Frame Attention (IAFA) mechanism for feature augmentation in the spatio-temporal domain. A dynamic fusion strategy helps refine forgery representations for the fusion. Moreover, we construct a comprehensive diffusion video dataset, called Diffusion Video Forensics (DVF), across a wide range of forgery videos. MM-Det achieves state-of-the-art performance in DVF, demonstrating the effectiveness of our algorithm. Both source code and DVF are available at https://github.com/SparkleXFantasy/MM-Det.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db63d17c-ece9-4c1a-8886-cb9a3ddfa65aCited by top-tier papers15
- Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionShuhai Zhang, Zihao Lian, Jiahao Yang, Daiyuan Li et al.NeurIPS 2025 · 29 citations
- Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented AugmentationRiccardo Corvi, Davide Cozzolino, Ekta Prashnani, Shalini De Mello et al.NeurIPS 2025 · 25 citations
- Tracing Hyperparameter Dependencies for Model Parsing via Learnable Graph Pooling NetworkXiao Guo, Vishal Asnani, Sijia Liu, Xiaoming LiuNeurIPS 2024 · 13 citations
- VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningHao Tan, jun lan, Senyuan Shi, Zichang Tan et al.ICML 2026 · 12 citations
- Mixture-of-Attack-Experts with Class Regularization for Unified Physical-Digital Face Attack DetectionShunxin Chen, Ajian Liu, Junze Zheng, Jun Wan et al.AAAI 2025 · 11 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake DetectionYoungseo Kim, Kwan Yun, Seokhyeon Hong, Sihun Cha et al.CVPR 2026 · 2 citations
- VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language ModelsXinan He, Yue Zhou, Bing Fan, Bin Li et al.NeurIPS 2025 · 20 citations
- Diffusion Facial Forgery DetectionHarry Cheng, Yangyang Guo, Tianyi Wang, Liqiang Nie et al.ACM MM 2024 · 42 citations
- GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation DetectionXiaocan Chen, Qilin Yin, Jiarui Liu, Wei Lu et al.AAAI 2025 · 3 citations
- ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in VideosPeijun Bao, Anwei Luo, Gang Pan, Alex C. Kot et al.CVPR 2026 · 2 citations
