ResProto-FD: Visual-Language Residual Prototype Sets for Generalized Face Forgery Detection
Jiuyao Jing, Yu Zheng, Chunlei Peng
Abstract
With the rapid development of generative models, such as generative adversarial networks and diffusion models, the task of face forgery detection has emerged, aiming to identify forged faces in real-world scenarios. A key challenge for current face forgery detection models is improving generalization to unknown forgeries. To address this, we propose ResProto-FD, a framework that constructs residual prototype sets to capture diverse forgery cues and discriminative differences from real faces. Our novel perspective collects prototypes from the most informative residual features generated during training, enabling better representation of various forgery traces and real-vs-fake distinctions. First, we introduce a Visual-Language Residual Learning (VLRL) module based on the CLIP model. This module constructs residual features between image and text embeddings to capture inconsistencies between visual features and associated textual semantics. In doing so, it guides the model to attend to subtle visual forgery clues and enhances the discriminative power of image representations. Furthermore, we design a Gradient-aware Residual Prototypes (GRP) mechanism— a dynamic collection strategy that selectively stores uncertain residual features based on gradient signals to build the prototype sets. This enhances the model’s ability to generalize to unknown forgery types. Extensive experiments across various datasets and forgery methods demonstrate that ResProto-FD significantly improves generalization performance and consistently outperforms state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Learning Self-Consistency for Deepfake DetectionTianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding et al.ICCV 2021 · 368 citations
- Detecting Deepfakes with Self-Blended ImagesKaede Shiohara, Toshihiko YamasakiCVPR 2022 · 366 citations
- End-to-End Reconstruction-Classification Learning for Face Forgery DetectionJunyi Cao, Chao Ma, Taiping Yao, Shen Chen et al.CVPR 2022 · 327 citations
Related papers
- HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery DetectionJialei Cui, Jianwei Du, Yanzhe Li, Lei Gao et al.ACM MM 2025 · 2 citations
- Unleashing Vision-Language Semantics for Deepfake Video DetectionJiawen Zhu, Yunqi Miao, Xueyi Zhang, Jiankang Deng et al.CVPR 2026
- Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorXiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu et al.CVPR 2025
- Forensics Adapter: Adapting CLIP for Generalizable Face Forgery DetectionXinjie Cui, Yuezun Li, Ao Luo, Jiaran Zhou et al.CVPR 2025
- DySy-Det: A Synergistic Framework with Dynamic Reconstruction-Path Consistency for AI-Generated Image DetectionFanli Jin, Feng Lin, Gaojian Wang, Tong Wu et al.AAAI 2026
