Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
Jongwook Choi, Taehoon Kim, Yonghyun Jeong, Seungryul Baek, Jongwon Choi
Abstract
This paper presents a new approach for the detection of fake videos, based on the analysis of style latent vectors and their abnormal behavior in temporal changes in the generated videos. We discovered that the generated facial videos suffer from the temporal distinctiveness in the temporal changes of style latent vectors, which are inevitable during the generation of temporally stable videos with various facial expressions and geometric transformations. Our framework utilizes the StyleGRU module, trained by contrastive learning, to represent the dynamic properties of style latent vectors. Additionally, we introduce a style attention module that integrates StyleGRU-generated features with contentbased features, enabling the detection of visual and temporal artifacts. We demonstrate our approach across various benchmark scenarios in deepfake detection, showing its superiority in cross-dataset and cross-manipulation scenarios. Through further analysis, we also validate the importance of using temporal changes of style latent vectors to improve the generality of deepfake video detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- FreqBlender: Enhancing DeepFake Detection by Blending Frequency KnowledgeHanzhe Li, Jiaran Zhou, Yuezun Li, Baoyuan Wu et al.NeurIPS 2024 · 96 citations
- X2-DFD: A framework for explainable and extendable Deepfake DetectionYize Chen, Zhiyuan Yan, Guangliang Cheng, Kangran Zhao et al.NeurIPS 2025 · 43 citations
- Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented AugmentationRiccardo Corvi, Davide Cozzolino, Ekta Prashnani, Shalini De Mello et al.NeurIPS 2025 · 25 citations
- From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face DeepfakesLong Ma, Zhiyuan Yan, Jin Xu, Yize Chen et al.NeurIPS 2025 · 24 citations
- Fair Deepfake Detectors Can GeneralizeHarry Cheng, Ming-Hui Liu, Yangyang Guo, Tianyi Wang et al.NeurIPS 2025 · 11 citations
Builds on33
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- ReStyle: A Residual-Based StyleGAN Encoder via Iterative RefinementYuval Alaluf, Or Patashnik, Daniel Cohen-OrICCV 2021 · 377 citations
Related papers
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh et al.ICCV 2025 · 7 citations
- Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and LocalizationKomal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan SubramanianACM MM 2020 · 217 citations
- Spatiotemporal Inconsistency Learning for DeepFake Video DetectionZhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding et al.ACM MM 2021 · 175 citations
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake DetectionYoungseo Kim, Kwan Yun, Seokhyeon Hong, Sihun Cha et al.CVPR 2026 · 2 citations
- Deepfake Video Detection with Spatiotemporal Dropout TransformerDaichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang et al.ACM MM 2022 · 46 citations
