AI-Generated Video Detection via Perceptual Straightening
Christian Internò, Robert Geirhos, Markus Olhofer, Sunny Liu, Barbara Hammer, David A. Klindt
摘要
The rapid advancement of generative AI enables highly realistic synthetic videos, posing significant challenges for content authentication and raising urgent concerns about misuse. Existing detection methods often struggle with generalization and capturing subtle temporal inconsistencies. We propose ReStraV(Representation Straightening for Video), a novel approach to distinguish natural from AI-generated videos. Inspired by the "perceptual straightening" hypothesis [1, 2]-which suggests real-world video trajectories become more straight in neural representation domain-we analyze deviations from this expected geometric property. Using a pre-trained self-supervised vision transformer (DINOv2), we quantify the temporal curvature and stepwise distance in the model's representation domain. We aggregate statistics of these measures for each video and train a classifier. Our analysis shows that AI-generated videos exhibit significantly different curvature and distance patterns compared to real videos. A lightweight classifier achieves state-of-the-art detection performance (e.g., 97.17% accuracy and 98.63% AUROC on the VidProM benchmark [3]), substantially outperforming existing image-and video-based methods. ReStraV is computationally efficient, offering a low-cost and effective detection solution. This work provides new insights into using neural representation geometry for AI-generated video detection. Classifier (e.g., MLP) In representation SSE space, natural videos trace straighter paths than AIgenerated videos. The trajectory geometry provides a discriminative signal. SSE (e.g., DINOv2) Frames are processed by a SSE and we collect the embeddings. Classifier: AI-generated vs. natural Trajectories in representation domain AI-Generated AI-Generated vs. Natural Natural .. .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Skyra: AI-Generated Video Detection via Grounded Artifact ReasoningYifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun 等CVPR 2026 · 被引用 24 次
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero 等ICML 2026 · 被引用 19 次
- VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningHao Tan, jun lan, Senyuan Shi, Zichang Tan 等ICML 2026 · 被引用 12 次
- Training-free Detection of Generated Videos via Spatial-Temporal LikelihoodsOmer Ben Hayun, Roy Betser, Meir Yossef Levi, Levi Kassel 等CVPR 2026 · 被引用 7 次
- Explainable Forensics of Manipulated Segments in Untrimmed Long VideosYue Feng, Jingjing Li, Qijia Lu, Wei Ji 等ICML 2026
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionShuhai Zhang, Zihao Lian, Jiahao Yang, Daiyuan Li 等NeurIPS 2025 · 被引用 29 次
- D3: Training-Free AI-Generated Video Detection Using Second-Order FeaturesChende Zheng, Ruiqi Suo, Chenhao Lin, Zhengyu Zhao 等ICCV 2025 · 被引用 11 次
- Detecting Generated Images by Fitting Natural Image DistributionsYonggang Zhang, Jun Nie, Xinmei Tian, Mingming Gong 等NeurIPS 2025 · 被引用 9 次
- Preserving Forgery Artifacts: AI-Generated Video Detection at Native ScaleZhengcen Li, Chenyang Jiang, Hang Zhao, Shiyang Zhou 等ICLR 2026 · 被引用 8 次
- Learning predictable and robust neural representations by straightening image sequencesXueyan Niu, Cristina Savin, Eero P. SimoncelliNeurIPS 2024 · 被引用 13 次
