Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
Junhao Xu, Jingjing Chen, Yang Jiao, Jiacheng Zhang, Zhiyu Tan, Hao Li, Yu-Gang Jiang
Abstract
Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns about digital media authenticity. While existing detection methods demonstrate promising results on benchmark datasets, they face critical limitations in real-world applications. First, existing detectors typically fail to detect semantic inconsistencies with the person’s identity, such as implausible behaviors or incompatible environmental contexts in given images. Second, these methods rely heavily on low-level visual cues, making them effective for known forgeries but less reliable against new or unseen manipulation techniques. To address these challenges, we present a novel personalized vision-language model (VLM) that integrates low-level visual artifact analysis and high-level semantic inconsistency detection. Unlike previous VLM-based methods, our approach avoids resource-intensive supervised fine-tuning that often struggles to preserve distinct identity characteristics. Instead, we employ a lightweight method that dynamically encodes identity-specific information into specialized identifier tokens. This design enables the model to learn distinct identity characteristics while maintaining robust generalization capabilities. We further enhance detection capabilities through a lightweight detection adapter that extracts fine-grained information from shallow features of the vision encoder, preserving critical low-level evidence. Comprehensive experiments demonstrate that our approach achieves 94.25% accuracy and 94.08% F1 score, outperforming both traditional forgery detectors and general VLMs while requiring only 10 extra tokens.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- WildDeepfake: A Challenging Real-World Dataset for Deepfake DetectionBojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma et al.ACM MM 2020 · 443 citations
- SimSwap: An Efficient Framework For High Fidelity Face SwappingRenwang Chen, Xuanhong Chen, Bingbing Ni, Yanhao GeACM MM 2020 · 409 citations
- End-to-End Reconstruction-Classification Learning for Face Forgery DetectionJunyi Cao, Chao Ma, Taiping Yao, Shen Chen et al.CVPR 2022 · 327 citations
- Exploring Temporal Coherence for More General Video Face Forgery DetectionYinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng et al.ICCV 2021 · 314 citations
Related papers
- Unleashing Vision-Language Semantics for Deepfake Video DetectionJiawen Zhu, Yunqi Miao, Xueyi Zhang, Jiankang Deng et al.CVPR 2026
- ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural NetworksAhmad Albarqawi, Mahmoud Nazzal, Issa Khalil, Abdallah Khreishah et al.NDSS 2026 · 1 citation
- Guard Me If You Know Me: Protecting Specific Face-Identity from DeepfakesKaiqing Lin, Zhiyuan Yan, Ke-Yue Zhang, Li Hao et al.NeurIPS 2025 · 10 citations
- Learning Forgery-Aware Lip Representations Without Forgery PriorsBofan Chen, Hongyu Zhu, Yi He, Sichu Liang et al.CVPR 2026
- ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationQing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu et al.CVPR 2026 · 3 citations
