MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng
Abstract
The proliferation of face forgery content on the Web poses a severe threat to online trust, social media security, and the credibility of digital information. Existing detection approaches often fail to generalize across diverse forgery types and unseen scenarios commonly encountered in web-scale applications. Recent studies have utilized visual large language models (VLMs) to answer not only ''Is this face a forgery?'' but also ''Why is the face a forgery?'' These studies introduced forgery-related attributes, such as forgery location and type, to construct deepfake VQA datasets and train VLMs, achieving high accuracy while providing human-understandable explanatory text descriptions. However, these methods still have limitations. For example, they do not fully leverage face quality-related attributes, which are often abnormal in forged faces, and they lack effective training strategies for forgery-aware VLMs. In this paper, we extend the VQA dataset to create DD-VQA+, which features a richer set of attributes and a more diverse range of samples. Furthermore, we introduce a novel forgery detection framework, MGFFD-VLM, which integrates an Attribute-Driven Hybrid LoRA Strategy to enhance the capabilities of Visual Large Language Models (VLMs). Additionally, our framework incorporates Multi-Granularity Prompt Learning and a Forgery-Aware Training Strategy. By transforming classification and forgery segmentation results into prompts, our method not only improves forgery classification but also enhances interpretability. To further boost detection performance, we design multiple forgery-related auxiliary losses. Experimental results demonstrate that our approach surpasses existing methods in both text-based forgery judgment and analysis, achieving superior accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fe03269-3e74-4894-8a4d-6bec53570c4eCited by top-tier papers1
Ask how each one uses itBuilds on7
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- WildDeepfake: A Challenging Real-World Dataset for Deepfake DetectionBojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma et al.ACM MM 2020 · 443 citations
- End-to-End Reconstruction-Classification Learning for Face Forgery DetectionJunyi Cao, Chao Ma, Taiping Yao, Shen Chen et al.CVPR 2022 · 327 citations
- CLIB-FIQA: Face Image Quality Assessment with Confidence CalibrationFu-Zhao Ou, Chongyi Li, Shiqi Wang, Sam KwongCVPR 2024 · 24 citations
Related papers
- Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake DetectionPeipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng et al.ICML 2025
- VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language ModelsXinan He, Yue Zhou, Bing Fan, Bin Li et al.NeurIPS 2025 · 20 citations
- Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorXiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu et al.CVPR 2025
- Towards General Visual-Linguistic Face Forgery DetectionKe Sun, Shen Chen, Taiping Yao, Ziyin Zhou et al.CVPR 2025
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
