Deepfake Video Detection with Spatiotemporal Dropout Transformer
Daichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang, Dan Zeng, Shiming Ge
摘要
While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each frame. Existing image-level approaches often focus on single frame and ignore the spatiotemporal cues hidden in deepfake videos, resulting in poor generalization and robustness. The key of a video-level detector is to fully exploit the spatiotemporal inconsistency distributed in local facial regions across different frames in deepfake videos. Inspired by that, this paper proposes a simple yet effective patch-level approach to facilitate deepfake video detection via spatiotemporal dropout transformer. The approach reorganizes each input video into bag of patches that is then fed into a vision transformer to achieve robust representation. Specifically, a spatiotemporal dropout operation is proposed to fully explore patch-level spatiotemporal cues and serve as effective data augmentation to further enhance model's robustness and generalization ability. The operation is flexible and can be easily plugged into existing vision transformers. Extensive experiments demonstrate the effectiveness of our approach against 25 state-of-the-arts with impressive robustness, generalizability, and representation ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu 等ACM MM 2023 · 被引用 42 次
- Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based CustomizationYisu Liu, Jinyang An, Wanqian Zhang, Dayan Wu 等ACM MM 2024 · 被引用 16 次
- RetouchingFFHQ: A Large-scale Dataset for Fine-grained Face Retouching DetectionQichao Ying, Jiaxin Liu, Sheng Li, Haisheng Xu 等ACM MM 2023 · 被引用 13 次
- SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake DetectionInzamamul Alam, Md Tanvir Islam, Simon S. WooACM MM 2025 · 被引用 6 次
- Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video DetectionDat Nguyen, Marcella Astrid, Anis Kacem, Enjie Ghorbel 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh 等ICCV 2025 · 被引用 7 次
- Delving into Sequential Patches for Deepfake DetectionJiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding 等NeurIPS 2022 · 被引用 84 次
- Face Forgery Detection via Symmetric TransformerLuchuan Song, Xiaodan Li, Zheng Fang, Zhenchao Jin 等ACM MM 2022 · 被引用 20 次
- Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter TuningZhiyuan Yan, Yandan Zhao, Shen Chen, Mingyi Guo 等CVPR 2025
- Protecting Celebrities from DeepFake with Identity Consistency TransformerXiaoyi Dong, Jianmin Bao, Dongdong Chen, Ting Zhang 等CVPR 2022 · 被引用 155 次
