VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
Kyoungjun Park, Yifan Yang, Juheon Yi, Shicheng Zheng, Muhammad Muaz, Yifei Shen, Dongqi Han, Caihua Shan, Lili Qiu
Abstract
The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based detectors rely on supervised fine-tuning (SFT) or direct preference optimization (DPO), these methods are often bottlenecked by static, pre-labeled datasets that fail to capture the evolving, multi-step physical inconsistencies of modern generative models. To bridge this gap, we introduce VidGuard-R1, the first video authenticity detector to utilize group relative policy optimization (GRPO). Moving beyond passive preference matching, VidGuard-R1 employs a reinforcement learning framework that encourages the model to explore and rank multiple reasoning paths. By introducing specialized reward models for temporal stability and diffusion-aware complexity, we incentivize the model to discover 'physics-grounded' artifacts. Our contributions include: (1) a curated dataset of 140,000 challenging real/fake video pairs; (2) a GRPO-based training paradigm that achieves state-of-the-art zero-shot performance; and (3) a reasoning-first architecture that provides precise, verifiable rationales for its forensic judgments. Project website: https://vidguard-r1.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7bbec31-e8a0-4dfd-80c3-e57afee4200eCited by top-tier papers4
- Skyra: AI-Generated Video Detection via Grounded Artifact ReasoningYifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun et al.CVPR 2026 · 24 citations
- VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningHao Tan, jun lan, Senyuan Shi, Zichang Tan et al.ICML 2026 · 12 citations
- COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsPeizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou et al.CVPR 2026 · 1 citation
- Geometric Flow Grounding: A Unified Manifold Decoupling Framework for Dynamics Discovery and VerificationChang Yu, Yuxuan Luo, Yixuan Du, Yuqing Zhou et al.ICML 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
Related papers
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake DetectionTuan Nguyen, Naseem Khan, Khang Tran, Hai Phan et al.ICML 2026
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPOJinyoung Park, Jeehye Na, Jinyoung Kim, Hyunwoo J. KimNeurIPS 2025 · 64 citations
- Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video DetectionShuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong et al.ICML 2026
- RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake DetectionTianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He et al.ACM MM 2025 · 3 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
