Spatio-Temporal Catcher: A Self-Supervised Transformer for Deepfake Video Detection
Maosen Li, Xurong Li, Kun Yu, Cheng Deng, Heng Huang, Feng Mao, Hui Xue, Minghao Li
Abstract
As deepfake technology has become increasingly sophisticated and accessible, making it easier for individuals with malicious intent to create convincing fake content, which has raised considerable concern in the multimedia and computer vision community. Despite significant advances in deepfake video detection, most existing methods mainly focused on model architecture and training processes with little focus on data perspectives. In this paper, we argue that data quality has become the main bottleneck of current research. To be specific, in the pre-training phase, the domain shift between pre-training and target datasets may lead to poor generalization ability. Meanwhile, in the training phase, the low fidelity of the existing datasets leads to detectors relying on specific low-level visual artifacts or inconsistency. To overcome the shortcomings, (1). In the pre-training phase, pre-train our model on high-quality facial videos by utilizing data-efficient reconstruction-based self-supervised learning to solve domain shift. (2). In the training phase, we develop a novel spatio-temporal generator that can synthesize various high-quality "fake" videos in large quantities at a low cost, which enables our model to learn more general spatio-temporal representations in a self-supervised manner. (3). Additinally, to take full advantage of synthetic "fake" videos, we adopt diversity losses at both frame and video levels to explore the diversity of clues in "fake" videos. Our proposed framework is data-efficient and does not require any real-world deepfake videos. Extensive experiments demonstrate that our method significantly improves the generalization capability. Particularly on the most challenging CDF and DFDC datasets, our method outperforms the baselines by 8.88% and 7.73% points, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f8a1c807-c0dd-421d-a304-3a95e194253bCited by top-tier papers2
- DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery AnalysisYinqi Cai, Jichang Li, Zhaolun Li, Weikai Chen et al.ICCV 2025 · 11 citations
- Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video DetectionDat Nguyen, Marcella Astrid, Anis Kacem, Enjie Ghorbel et al.ICCV 2025 · 5 citations
Related papers
- Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation ModelYue-Hua Han, Tai-Ming Huang, Kai-Lung Hua, Jun-Cheng ChenCVPR 2025
- Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter TuningZhiyuan Yan, Yandan Zhao, Shen Chen, Mingyi Guo et al.CVPR 2025
- Improving the Efficiency and Robustness of Deepfakes Detection Through Precise Geometric FeaturesZekun Sun, Yujie Han, Zeyu Hua, Na Ruan et al.CVPR 2021
- FakeDiffer: Distributional Disparity Learning on Differentiated Reconstruction for Face Forgery DetectionBo Wang, Zhao Zhang, Suiyi Zhao, Xianming Ye et al.AAAI 2025 · 4 citations
- Celeb-DF: A Large-Scale Challenging Dataset for DeepFake ForensicsYuezun Li, Xin Yang, Pu Sun, Honggang Qi et al.CVPR 2020
