Community Forensics: Using Thousands of Generators to Train Fake Image Detectors
Jeongsoo Park, Andrew Owens
Abstract
One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this problem, and we propose a new dataset that is significantly larger and more diverse than prior works. As part of creating this dataset, we systematically download thousands of text-to-image latent diffusion models and sample images from them. We also collect images from dozens of popular open source and commercial models. The resulting dataset contains 2.7M images that have been sampled from 4803 different models. These images collectively capture a wide range of scene content, generator architectures, and image processing settings. Using this dataset, we study the generalization abilities of fake image detectors. Our experiments suggest that detection performance improves as the number of models in the training set increases, even when these models have similar architectures. We also find that increasing the diversity of the models improves detection performance, and that our trained detectors generalize better than those trained on other datasets. The dataset can be found in
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 556a6a2c-a6c4-4098-9231-d79397b8edc2Cited by top-tier papers17
- Scaling Up AI-Generated Image Detection with Generator-Aware PrototypesZiheng Qin, Yuheng Ji, Renshuai Tao, Yuxuan Tian et al.CVPR 2026 · 10 citations
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning SupervisionKartik Kuckreja, Parul Gupta, Muhammad Haris Khan, Abhinav DhallCVPR 2026 · 5 citations
- Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake DetectionTianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He et al.CVPR 2026 · 5 citations
- Towards Reliable Identification of Diffusion-based Image ManipulationsAlex Costanzino, Woody Bayliss, Juil Sock, Marc Górriz Blanch et al.NeurIPS 2025 · 4 citations
Builds on54
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionYan Hong, Jianming Feng, Haoxing Chen, Jun Lan et al.AAAI 2025 · 13 citations
- RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated ImagesHanzhe Yu, Yun Ye, Jintao Rong, Qi Xuan et al.ACM MM 2025 · 1 citation
- AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness BenchmarkLi Lin, Santosh Santosh, Mingyang Wu, Xin Wang et al.CVPR 2025
- FakeInversion: Learning to Detect Images from Unseen Text-to-Image Models by Inverting Stable DiffusionGeorge Cazenavette, Avneesh Sud, Thomas Leung, Ben UsmanCVPR 2024
- Breaking Semantic Artifacts for Generalized AI-generated Image DetectionChende Zheng, Chenhao Lin, Zhengyu Zhao, Hang Wang et al.NeurIPS 2024 · 57 citations
