Learning Forgery-Aware Lip Representations Without Forgery Priors
Bofan Chen, Hongyu Zhu, Yi He, Sichu Liang, Shi-Lin Wang
摘要
Visual Speaker Authentication (VSA) verifies identity by analyzing lip dynamics during prompted speech, offering enhanced privacy compared to full-face methods while maintaining discriminability for high-security applications. However, recent advances in personalized talking face generation (TFG) have enabled realistic forgeries that closely mimic lip dynamics in sync with speech, posing severe threats to VSA systems. Prevailing defenses rely heavily on supervised classifiers trained on known forgeries via empirical risk minimization, resulting in poor generalization to unseen attacks, dependency on continuously updated fake data, and complete failure in the absence of effective forgery priors. In this paper, we revisit the design of forgery detectors and argue that over-reliance on fake priors hinders the exploitation of rich authenticity signals inherently present in real videos. We propose a novel detector trained exclusively on authentic data, learning forgery-aware representations through three key components: (1) lightweight modules that capture forgery-indicative statistics from real videos; (2) an asymmetric contrastive objective that compacts real samples while repelling potential forgeries in representation space; and (3) a theoretically grounded regularizer that shapes real representations into a tractable, isotropic Gaussian. To support rigorous evaluation, we introduce a benchmark suite spanning diverse TFG forgeries. Across eight modern forgery attacks and ten state-of-the-art (SOTA) detectors, we achieve over a 10% reduction in error rates while preserving identity-verification capability with minimal overhead, and demonstrate robust generalization under diverse and complex real-world conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper38
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 被引用 755 次
相关 Paper
- Enhancing the Security of Visual Speaker Authentication Based on Dynamic Lip-Print AnalysisYi He, Lei Yang, Bofan Chen, Shilin WangCVPR 2026
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery DetectionYachao Liang, Min Yu, Gang Li, Jianguo Jiang 等NeurIPS 2024 · 被引用 19 次
- Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head PosesTianyi She, Jiawei Liu, Weifeng Liu, Hanqing Zhao 等ICML 2026
- Leveraging Real Talking Faces via Self-Supervision for Robust Forgery DetectionAlexandros Haliassos, Rodrigo Mira, Stavros Petridis, Maja PanticCVPR 2022 · 被引用 138 次
- Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery DetectionAlexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, Maja PanticCVPR 2021
