FiSeR: Fine-Grained Source Representations for Cross-Domain AI Image Detection
Shan Zhang, Yongxin He, Mingming Zhang, Huiwen Tian, Lei Ma
Abstract
Real-world synthetic image detectors often generalize poorly under domain shift despite strong indomain performance. Using unsupervised UMAP projections, we find that natural and synthetic features remain partially separable on unseen datasets, yet performance still drops, suggesting that the classification head overfits to trainingdomain artifacts. Therefore, the key is to learn more transferable representations so that the decision criterion is more stable and robust to domain shifts. Based on the structural fact that synthetic images are produced by diverse generators, we propose a hierarchical contrastive learning framework that improves the separability between natural and synthetic images while preserving generator identity information. It jointly optimizes (i) a coarse contrastive objective between natural and synthetic images and (ii) a fine contrastive objective among synthetic images using generator identities. Trained on Wild-Fake, our method achieves an average AUROC gain of +10.22 on cross-domain evaluation over Chameleon, AIGIBench, Community Forensics, and GenImage under the same settings as the strong baseline DIRE. For few-shot adaptation, we freeze the backbone and fit an SVM head on 10 labeled samples per class, improving AU-ROC by +10.64 on AIGIBench and +17.41 on Chameleon, averaged over 12 widely used detectors. Our code is publicly available at: https: //github.com/heyongxin233/FiSeR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- DIRE for Diffusion-Generated Image DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang et al.ICCV 2023 · 479 citations
- Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningChuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu et al.AAAI 2024 · 232 citations
Related papers
- OA-FSUI2IT: A Novel Few-Shot Cross Domain Object Detection Framework with Object-Aware Few-Shot Unsupervised Image-to-Image TranslationLifan Zhao, Yunlong Meng, Lin XuAAAI 2022 · 4 citations
- One for All: Synthesis-Free Fingerprint Learning for Attribution of In-the-Wild Synthetic ImagesJianwei Fei, Yunshu Dai, Peipeng Yu, Zhihua Xia et al.AAAI 2026
- Aggregating Diverse Cue Experts for AI-Generated Image DetectionLei Tan, Shuwei Li, Mohan Kankanhalli, Robby T. TanAAAI 2026
- ConFeSS: A Framework for Single Source Cross-Domain Few-Shot LearningDebasmit Das, Sungrack Yun, Fatih PorikliICLR 2022 · 57 citations
- SimLBR: Learning to Detect Fake Images by Learning to Detect Real ImagesAayush Dhakal, Subash Khanal, Srikumar Sastry, Jacob Arndt et al.CVPR 2026 · 1 citation
