Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector
Xiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu, Xiaoming Liu
摘要
Deepfake detection is a long-established research topic vital for mitigating the spread of malicious misinformation. Unlike prior methods that provide either binary classification results or textual explanations separately, we introduce a novel method capable of generating both simultaneously. Our method harnesses the multi-modal learning capability of the pre-trained CLIP and the unprecedented interpretability of large language models (LLMs) to enhance both the generalization and explainability of deepfake detection. Specifically, we introduce a multi-modal face forgery detector (M2F2-Det) that employs tailored face forgery prompt learning, incorporating the pre-trained CLIP to improve generalization to unseen forgeries. Also, M2F2-Det incorporates an LLM to provide detailed textual explanations of its detection decisions, enhancing interpretability by bridging the gap between natural language and subtle cues of facial forgeries. Empirically, we evaluate M2F2-Det on both detection and explanation generation tasks, where it achieves state-of-the-art performance, demonstrating its effectiveness in identifying and explaining diverse forgeries. Source code is available at link. Recently, the powerful capability of vision-language models, e.g., CLIP [56], also inspired efforts in detecting deepfakes. For example, DDVQA-BLIP [85] reformulates deepfake detection as an explanation generation task using a vision-language model [37], which enhances interpretability through natural language descriptions (Fig. 1b ). In addition, several binary detectors [10, 52, 60 ] leverage CLIP's robust recognition capabilities to achieve impressive performance. However, three key limitations remain in these works. First, DDVQA-BLIP relies on a general text-generation model without dedicated mechanisms for deepfake detection, resulting in lower detection accuracy compared to conventional binary detectors. Secondly, prior CLIP-based detectors often lack effective input text prompts to describe diverse forgeries, restricting the adaptation of CLIP's multi-modal learning ability in the detection task. Third, while CLIP's open-set recognition capability -enabling it to identify diverse visual semantics -is successfully combined with LLMs in domains like document parsing [25, 48, 79] and medical diagnosis [34, 50, 83], its integration with LLMs for deepfake detection remains largely unexplored.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsDingqiang Ye, Chao Fan, Zhanbo Huang, Chengwen Luo 等NeurIPS 2025 · 被引用 28 次
- Veritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningHao Tan, Jun Lan, Zichang Tan, Senyuan Shi 等ICLR 2026 · 被引用 26 次
- Skyra: AI-Generated Video Detection via Grounded Artifact ReasoningYifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun 等CVPR 2026 · 被引用 24 次
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain 等CVPR 2026 · 被引用 7 次
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated ImagesChuangchuang Tan, Xiang Ming, Jinglu Wang, Renshuai Tao 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake DetectionPeipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng 等ICML 2025
- Unleashing Vision-Language Semantics for Deepfake Video DetectionJiawen Zhu, Yunqi Miao, Xueyi Zhang, Jiankang Deng 等CVPR 2026
- Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake DetectionKaiqing Lin, Yuzhen Lin, Weixiang Li, Taiping Yao 等AAAI 2025 · 被引用 32 次
- X2-DFD: A framework for explainable and extendable Deepfake DetectionYize Chen, Zhiyuan Yan, Guangliang Cheng, Kangran Zhao 等NeurIPS 2025 · 被引用 43 次
- C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake DetectionChuangchuang Tan, Renshuai Tao, Huan Liu, Guanghua Gu 等AAAI 2025 · 被引用 92 次
