Deepfake Video Detection with Spatiotemporal Dropout Transformer
Daichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang, Dan Zeng, Shiming Ge
Abstract
While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each frame. Existing image-level approaches often focus on single frame and ignore the spatiotemporal cues hidden in deepfake videos, resulting in poor generalization and robustness. The key of a video-level detector is to fully exploit the spatiotemporal inconsistency distributed in local facial regions across different frames in deepfake videos. Inspired by that, this paper proposes a simple yet effective patch-level approach to facilitate deepfake video detection via spatiotemporal dropout transformer. The approach reorganizes each input video into bag of patches that is then fed into a vision transformer to achieve robust representation. Specifically, a spatiotemporal dropout operation is proposed to fully explore patch-level spatiotemporal cues and serve as effective data augmentation to further enhance model's robustness and generalization ability. The operation is flexible and can be easily plugged into existing vision transformers. Extensive experiments demonstrate the effectiveness of our approach against 25 state-of-the-arts with impressive robustness, generalizability, and representation ability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb853922-f248-4e93-93e7-cc1502e37896Cited by top-tier papers5
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu et al.ACM MM 2023 · 42 citations
- Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based CustomizationYisu Liu, Jinyang An, Wanqian Zhang, Dayan Wu et al.ACM MM 2024 · 16 citations
- RetouchingFFHQ: A Large-scale Dataset for Fine-grained Face Retouching DetectionQichao Ying, Jiaxin Liu, Sheng Li, Haisheng Xu et al.ACM MM 2023 · 13 citations
- SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake DetectionInzamamul Alam, Md Tanvir Islam, Simon S. WooACM MM 2025 · 6 citations
- Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video DetectionDat Nguyen, Marcella Astrid, Anis Kacem, Enjie Ghorbel et al.ICCV 2025 · 5 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh et al.ICCV 2025 · 7 citations
- Delving into Sequential Patches for Deepfake DetectionJiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding et al.NeurIPS 2022 · 84 citations
- Face Forgery Detection via Symmetric TransformerLuchuan Song, Xiaodan Li, Zheng Fang, Zhenchao Jin et al.ACM MM 2022 · 20 citations
- Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter TuningZhiyuan Yan, Yandan Zhao, Shen Chen, Mingyi Guo et al.CVPR 2025
- Protecting Celebrities from DeepFake with Identity Consistency TransformerXiaoyi Dong, Jianmin Bao, Dongdong Chen, Ting Zhang et al.CVPR 2022 · 155 citations
