Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
Shuhai Zhang, Zihao Lian, Jiahao Yang, Daiyuan Li, Guoxuan Pang, Feng Liu, Bo Han, Shutao Li, Mingkui Tan
摘要
AI-generated videos have achieved near-perfect visual realism (e.g., Sora), urgently necessitating reliable detection mechanisms. However, detecting such videos faces significant challenges in modeling high-dimensional spatiotemporal dynamics and identifying subtle anomalies that violate physical laws. In this paper, we propose the first physics-driven AI-generated video detection paradigm based on probability flow conservation principles. Specifically, we propose a statistic called Normalized Spatiotemporal Gradient (NSG), which quantifies the ratio of spatial probability gradients to temporal density changes, explicitly capturing deviations from natural video dynamics. Leveraging pre-trained diffusion models, we develop an NSG estimator through spatial gradients approximation and motion-aware temporal modeling without complex motion decomposition while preserving physical constraints. Building on this, we propose an NSG-based video detection method (NSG-VD) that computes the Maximum Mean Discrepancy (MMD) between NSG features of the test and real videos as a detection metric. Last, we derive an upper bound of NSG feature distances between real and generated videos, proving that generated videos exhibit amplified discrepancies due to distributional shifts. Extensive experiments confirm that NSG-VD outperforms state-of-the-art baselines by 16.00% in Recall and 10.75% in F1-Score, validating the superior performance of NSG-VD. The source code is available at https://github.com/ZSHsh98/NSG-VD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Skyra: AI-Generated Video Detection via Grounded Artifact ReasoningYifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun 等CVPR 2026 · 被引用 24 次
- VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningHao Tan, jun lan, Senyuan Shi, Zichang Tan 等ICML 2026 · 被引用 12 次
- Training-free Detection of Generated Videos via Spatial-Temporal LikelihoodsOmer Ben Hayun, Roy Betser, Meir Yossef Levi, Levi Kassel 等CVPR 2026 · 被引用 7 次
- Explainable Forensics of Manipulated Segments in Untrimmed Long VideosYue Feng, Jingjing Li, Qijia Lu, Wei Ji 等ICML 2026
- Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video DetectionShuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong 等ICML 2026
它引用的顶会 Paper44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 被引用 1,527 次
相关 Paper
- D3: Training-Free AI-Generated Video Detection Using Second-Order FeaturesChende Zheng, Ruiqi Suo, Chenhao Lin, Zhengyu Zhao 等ICCV 2025 · 被引用 11 次
- Physical Simulator In-the-Loop Video GenerationLin Geng Foo, Mark He Huang, Alexandros Lattas, Stylianos Moschoglou 等CVPR 2026 · 被引用 13 次
- AI-Generated Video Detection via Perceptual StraighteningChristian Internò, Robert Geirhos, Markus Olhofer, Sunny Liu 等NeurIPS 2025 · 被引用 44 次
- How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation ApproachChirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu 等ICCV 2025 · 被引用 15 次
- NS-Diff: Fluid Navier-Stokes Guided Video Diffusion via Reinforcement LearningZijun Deng, Yuxin PengCVPR 2026
