Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
Zhiyuan Yan, Yandan Zhao, Shen Chen, Mingyi Guo, Xinghe Fu, Taiping Yao, Shouhong Ding, Yunsheng Wu, Li Yuan
Abstract
Three key challenges hinder the development of current deepfake video detection: (1) Temporal features can be complex and diverse: how can we identify general temporal artifacts to enhance model generalization? (2) Spatiotemporal models are proven to lean heavily on one type of forgery artifact and ignore the other (e.g., learning spatial only): how can we ensure balanced learning from both? (3) Videos are naturally resource-intensive: how can we tackle efficiency without compromising accuracy? This paper attempts to tackle the three challenges jointly. First, inspired by the notable generality of using image-level blending data for image forgery detection, we investigate whether and how video-level blending can be effective in video. We then perform a thorough analysis and identify a previously underexplored temporal forgery artifact: Facial Feature Drift (FFD), which commonly exists across different deepfakes. To reproduce FFD, we then propose a novel Videolevel Blending data (VB), which is implemented by blending the original image and its warped version frame-byframe, serving as a hard negative sample to mine more general artifacts. Second, we carefully design a lightweight Spatiotemporal Adapter (StA) to equip a pretrained image model with the ability to capture both spatial and temporal features jointly and efficiently. StA is designed with twostream 3D-Conv with varying kernel sizes, allowing it to process spatial and temporal features separately. Extensive experiments validate the effectiveness of our methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c195eb6-a51f-4102-a1d7-7b0e465fe6c9Cited by top-tier papers18
- X2-DFD: A framework for explainable and extendable Deepfake DetectionYize Chen, Zhiyuan Yan, Guangliang Cheng, Kangran Zhao et al.NeurIPS 2025 · 43 citations
- Veritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningHao Tan, Jun Lan, Zichang Tan, Senyuan Shi et al.ICLR 2026 · 26 citations
- Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented AugmentationRiccardo Corvi, Davide Cozzolino, Ekta Prashnani, Shalini De Mello et al.NeurIPS 2025 · 25 citations
- From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face DeepfakesLong Ma, Zhiyuan Yan, Jin Xu, Yize Chen et al.NeurIPS 2025 · 24 citations
- VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningHao Tan, jun lan, Senyuan Shi, Zichang Tan et al.ICML 2026 · 12 citations
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Deepfake Video Detection with Spatiotemporal Dropout TransformerDaichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang et al.ACM MM 2022 · 46 citations
- AltFreezing for More General Video Face Forgery DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang et al.CVPR 2023
- Spatio-Temporal Catcher: A Self-Supervised Transformer for Deepfake Video DetectionMaosen Li, Xurong Li, Kun Yu, Cheng Deng et al.ACM MM 2023 · 9 citations
- Exploiting Style Latent Flows for Generalizing Deepfake Video DetectionJongwook Choi, Taehoon Kim, Yonghyun Jeong, Seungryul Baek et al.CVPR 2024
- Delving into Sequential Patches for Deepfake DetectionJiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding et al.NeurIPS 2022 · 84 citations
