Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
Peipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng, Zhihua Xia, Chip-Hong Chang
Abstract
Current Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs' potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-ofthe-art methods in generalization performance, while also supporting multi-turn dialogue capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35148515-78a7-412f-a66d-6fc0eeb2018aCited by top-tier papers5
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake DetectionJian-Yu Jiang-Lin, Kang-Yang Huang, Ling Zou, Ling Lo et al.CVPR 2026 · 5 citations
- Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsZiyin Zhou, Yunpeng Luo, Yuanchen Wu, Ke Sun et al.ICCV 2025 · 3 citations
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness AsymmetryXinchang Wang, Yunhao Chen, Yuechen Zhang, Congcong Bian et al.ICML 2026 · 2 citations
- Fine-Grained DINO Tuning with Dual Supervision for Face Forgery DetectionTianxiang Zhang, Peipeng Yu, Zhihua Xia, Longchen Dai et al.AAAI 2026 · 1 citation
- One for All: Synthesis-Free Fingerprint Learning for Attribution of In-the-Wild Synthetic ImagesJianwei Fei, Yunshu Dai, Peipeng Yu, Zhihua Xia et al.AAAI 2026
Builds on33
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorXiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu et al.CVPR 2025
- MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLMTao Chen, Jingyi Zhang, Decheng Liu, Chunlei PengWWW 2026 · 1 citation
- VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language ModelsXinan He, Yue Zhou, Bing Fan, Bin Li et al.NeurIPS 2025 · 20 citations
- Unleashing Vision-Language Semantics for Deepfake Video DetectionJiawen Zhu, Yunqi Miao, Xueyi Zhang, Jiankang Deng et al.CVPR 2026
- Towards General Visual-Linguistic Face Forgery DetectionKe Sun, Shen Chen, Taiping Yao, Ziyin Zhou et al.CVPR 2025
