Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses
Tianyi She, Jiawei Liu, Weifeng Liu, Hanqing Zhao, Weiming Zhang, Kejiang Chen
摘要
Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as state-of-the-art generative models not only optimize for the lip synchronization but also significantly eliminate visual artifacts, resulting in the lack of key detection signals. Inspired by the inherent biological coupling between lip movements and head poses in natural speech, we observe that generative models fundamentally disrupt this global coordination when optimizing for local lip motion. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as generative fingerprints, enabling source tracing. To validate our approach, we conduct extensive experiments on two challenging LipSync benchmarks as well as on our own proposed large-scale and multi-generator dataset, LipSyncBench-A. LipDA achieves over 97% AUC in detection and 97.5% accuracy in model attribution, significantly outperforming existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Exploring Temporal Coherence for More General Video Face Forgery DetectionYinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng 等ICCV 2021 · 被引用 314 次
- Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningChuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu 等AAAI 2024 · 被引用 232 次
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 被引用 219 次
相关 Paper
- Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakesWeifeng Liu, Tianyi She, Jiawei Liu, Boheng Li 等NeurIPS 2024 · 被引用 57 次
- Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery DetectionAlexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, Maja PanticCVPR 2021
- Learning Forgery-Aware Lip Representations Without Forgery PriorsBofan Chen, Hongyu Zhu, Yi He, Sichu Liang 等CVPR 2026
- From Talking to Singing: A New Challenge for Audio-Visual Deepfake DetectionKe Liu, Jiwei Wei, Wenyu Zhang, Shuchang Zhou 等ICML 2026
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery DetectionYachao Liang, Min Yu, Gang Li, Jianguo Jiang 等NeurIPS 2024 · 被引用 19 次
