Delving into Sequential Patches for Deepfake Detection
Jiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding, Jingdong Wang, Chengbin Quan, Youjian Zhao
Abstract
Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal information in pursuit to generalize well across deepfake methods, however, they still suffer from robustness problem against post-processings. In this work, we propose the Local-& Temporal-aware Transformer-based Deepfake Detection (LTTD) framework, which adopts a local-to-global learning protocol with a particular focus on the valuable temporal information within local sequences. Specifically, we propose a Local Sequence Transformer (LST), which models the temporal consistency on sequences of restricted spatial regions, where low-level information is hierarchically enhanced with shallow layers of learned 3D filters. Based on the local temporal embeddings, we then achieve the final classification in a global contrastive way. Extensive experiments on popular datasets validate that our approach effectively spots local forgery cues and achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dddcd2d9-5fec-44b4-9cd7-1fca8efd6d48Cited by top-tier papers8
- Can We Leave Deepfake Data Behind in Training Deepfake Detector?Jikang Cheng, Zhiyuan Yan, Ying Zhang, Yuhao Luo et al.NeurIPS 2024 · 85 citations
- SeeABLE: Soft Discrepancies and Bounded Contrastive Learning for Exposing DeepfakesNicolas Larue, Ngoc-Son Vu, Vitomir Struc, Peter Peer et al.ICCV 2023 · 58 citations
- Locate and Verify: A Two-Stream Network for Improved Deepfake DetectionChao Shuai, Jieming Zhong, Shuang Wu, Feng Lin et al.ACM MM 2023 · 52 citations
- Adversarial Robust Safeguard for Evading Deep Facial ManipulationJiazhi Guan, Yi Zhao, Zhuoer Xu, Changhua Meng et al.AAAI 2024 · 11 citations
- Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object DetectorsMingqian Ji, Shanshan Zhang, Jian YangCVPR 2026 · 1 citation
Builds on32
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- Deepfake Video Detection with Spatiotemporal Dropout TransformerDaichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang et al.ACM MM 2022 · 46 citations
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh et al.ICCV 2025 · 7 citations
- Face Forgery Detection via Symmetric TransformerLuchuan Song, Xiaodan Li, Zheng Fang, Zhenchao Jin et al.ACM MM 2022 · 20 citations
- Spatiotemporal Inconsistency Learning for DeepFake Video DetectionZhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding et al.ACM MM 2021 · 175 citations
- Exploring Temporal Coherence for More General Video Face Forgery DetectionYinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng et al.ICCV 2021 · 314 citations
