Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection
Huan Liu, Zichang Tan, Chuangchuang Tan, Yunchao Wei, Jingdong Wang, Yao Zhao
Abstract
In this paper, we study the problem of generalizable synthetic image detection, aiming to detect forgery images from diverse generative methods, e.g., GANs and diffusion models. Cutting-edge solutions start to explore the benefits of pre-trained models, and mainly follow the fixed paradigm of solely training an attached classifier, e.g., combining frozen CLIP-ViT with a learnable linear layer in UniFD [35]. However, our analysis shows that such a fixed paradigm is prone to yield detectors with insufficient learning regarding forgery representations. We attribute the key challenge to the lack of forgery adaptation, and present a novel forgeryaware adaptive transformer approach, namely FatFormer. Based on the pre-trained vision-language spaces of CLIP, FatFormer introduces two core designs for the adaption to build generalized forgery representations. First, motivated by the fact that both image and frequency analysis are essential for synthetic image detection, we develop a forgeryaware adapter to adapt image features to discern and integrate local forgery traces within image and frequency domains. Second, we find that considering the contrastive objectives between adapted image features and text prompt embeddings, a previously overlooked aspect, results in a nontrivial generalization improvement. Accordingly, we introduce language-guided alignment to supervise the forgery adaptation with image and text prompts in FatFormer. Experiments show that, by coupling these two designs, our approach tuned on 4-class ProGAN data attains a remarkable detection performance, achieving an average of 98% accuracy to unseen GANs, and surprisingly generalizes to unseen diffusion models with 95% accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a019f4b2-7147-4256-830c-b6962d8bd000Cited by top-tier papers70
- C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake DetectionChuangchuang Tan, Renshuai Tao, Huan Liu, Guanghua Gu et al.AAAI 2025 · 92 citations
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang et al.NeurIPS 2025 · 82 citations
- Dual Data Alignment Makes AI-Generated Image Detector Easier GeneralizableRuoxin Chen, Junwei Xi, Zhiyuan Yan, Ke-Yue Zhang et al.NeurIPS 2025 · 78 citations
- On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionXiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li et al.NeurIPS 2024 · 63 citations
- CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-SpoofingAjian Liu, Shuai Xue, Jianwen Gan, Jun Wan et al.CVPR 2024 · 59 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
Related papers
- DySy-Det: A Synergistic Framework with Dynamic Reconstruction-Path Consistency for AI-Generated Image DetectionFanli Jin, Feng Lin, Gaojian Wang, Tong Wu et al.AAAI 2026
- ForgeLens: Data-Efficient Forgery Focus for Generalizable Forgery Image DetectionYingjian Chen, Lei Zhang, Yakun NiuICCV 2025 · 2 citations
- Forensics Adapter: Adapting CLIP for Generalizable Face Forgery DetectionXinjie Cui, Yuezun Li, Ao Luo, Jiaran Zhou et al.CVPR 2025
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt LearningYiheng Li, Zichang Tan, Guoqing Xu, Zhen Lei et al.CVPR 2026 · 8 citations
- Scaling Up AI-Generated Image Detection with Generator-Aware PrototypesZiheng Qin, Yuheng Ji, Renshuai Tao, Yuxuan Tian et al.CVPR 2026 · 10 citations
