Delving into Sequential Patches for Deepfake Detection
Jiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding, Jingdong Wang, Chengbin Quan, Youjian Zhao
摘要
Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal information in pursuit to generalize well across deepfake methods, however, they still suffer from robustness problem against post-processings. In this work, we propose the Local-& Temporal-aware Transformer-based Deepfake Detection (LTTD) framework, which adopts a local-to-global learning protocol with a particular focus on the valuable temporal information within local sequences. Specifically, we propose a Local Sequence Transformer (LST), which models the temporal consistency on sequences of restricted spatial regions, where low-level information is hierarchically enhanced with shallow layers of learned 3D filters. Based on the local temporal embeddings, we then achieve the final classification in a global contrastive way. Extensive experiments on popular datasets validate that our approach effectively spots local forgery cues and achieves state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Can We Leave Deepfake Data Behind in Training Deepfake Detector?Jikang Cheng, Zhiyuan Yan, Ying Zhang, Yuhao Luo 等NeurIPS 2024 · 被引用 85 次
- SeeABLE: Soft Discrepancies and Bounded Contrastive Learning for Exposing DeepfakesNicolas Larue, Ngoc-Son Vu, Vitomir Struc, Peter Peer 等ICCV 2023 · 被引用 58 次
- Locate and Verify: A Two-Stream Network for Improved Deepfake DetectionChao Shuai, Jieming Zhong, Shuang Wu, Feng Lin 等ACM MM 2023 · 被引用 52 次
- Adversarial Robust Safeguard for Evading Deep Facial ManipulationJiazhi Guan, Yi Zhao, Zhuoer Xu, Changhua Meng 等AAAI 2024 · 被引用 11 次
- Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object DetectorsMingqian Ji, Shanshan Zhang, Jian YangCVPR 2026 · 被引用 1 次
它引用的顶会 Paper32
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- Deepfake Video Detection with Spatiotemporal Dropout TransformerDaichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang 等ACM MM 2022 · 被引用 46 次
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh 等ICCV 2025 · 被引用 7 次
- Face Forgery Detection via Symmetric TransformerLuchuan Song, Xiaodan Li, Zheng Fang, Zhenchao Jin 等ACM MM 2022 · 被引用 20 次
- Spatiotemporal Inconsistency Learning for DeepFake Video DetectionZhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding 等ACM MM 2021 · 被引用 175 次
- Exploring Temporal Coherence for More General Video Face Forgery DetectionYinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng 等ICCV 2021 · 被引用 314 次
